Audio ML Papers

Last 7 Days (August 07 - August 14, 2026)

Subcategories: All (17) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (16)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 74)
Feng Yin, Shuai Shi, Junjie Zheng ... ยท VUI Labs Research ยท arXiv
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Qu...
#2 TOP PAPER (Score: 73)
Haoyu Zhang, Zhipeng Li, Xiaoying Tang ... ยท The Chinese University of Hong Kong, Shenzhen +4 ยท arXiv
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and re...
#3 TOP PAPER (Score: 72)
Jiabao Zhuang, Changhao Jiang, Hanchen Wang ... ยท Fudan University ยท arXiv
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, an...
Thursday, August 13, 2026
Wenxiang Guo, Changhao Pan, Ziyue Jiang ... ยท Zhejiang University ยท IEEE Transactions on Multimedia
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing. Existing Text-to-Audio (T2A) systems either reduce quoted speech to unintellig...
Wednesday, August 12, 2026
Feng Yin, Shuai Shi, Junjie Zheng ... ยท VUI Labs Research ยท arXiv
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Qu...
Jiabao Zhuang, Changhao Jiang, Hanchen Wang ... ยท Fudan University ยท arXiv
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, an...
Yining Wang ยท arXiv
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence ...
Xingwei Sun, Heinrich Dinkel, Gang Li ... ยท Xiaomi Inc. +3 ยท arXiv
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal opti...
Rong Chao, Sung-Feng Huang, Moreno La Quatra, โ˜… Yu Tsao ... ยท National Taiwan University +4 ยท INTERSPEECH 2026
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidt...
Tuesday, August 11, 2026
Haoyu Zhang, Zhipeng Li, Xiaoying Tang ... ยท The Chinese University of Hong Kong, Shenzhen +4 ยท arXiv
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and re...
Gaopeng Xu, Zhenyu Wang, Zheng Xue ... ยท Qwen Business Unit of Alibaba ยท arXiv
The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to ...
Monday, August 10, 2026
Yan Rong, Fengji Ma, Xu Li ... ยท Kling Team ยท arXiv
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle to achieve these two goals and mainly rely on supervised fine-tuning, yielding sub-optimal performa...
Shuyu Li, Kejun Zhang, Jiahe Lei ... ยท Zhejiang University +3 ยท arXiv
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduc...
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah ... ยท ServiceNow ยท arXiv
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perc...
Zhanhong He, Hanyu Meng, David Defeng Huang ... ยท University of Western Australia +1 ยท ISMIR 2026
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper s...
Yanqiu Li, Yang Xiao, Jisheng Bai ... ยท arXiv
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet exist...
Saturday, August 08, 2026
Zixiang Wan, Xusheng Yang, Zheng Wang ... ยท Peking University +1 ยท arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J ... ยท Indian Institute of Science (IISc) +2 ยท arXiv
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and...
Friday, August 07, 2026
Shibo Wang, Zicheng Zhang, Libo Wang ... ยท Alibaba Group ยท arXiv
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call th...
Hanke Xie, Haopeng Lin, Jiale Qian, โ˜… Lei Xie ... ยท Shanghai Jiao Tong University +1 ยท arXiv
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit tok...