Audio ML Papers

Last 7 Days (August 05 - August 12, 2026)

Subcategories: All (11) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (10)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 73)
Haoyu Zhang, Zhipeng Li, Xiaoying Tang ... ยท The Chinese University of Hong Kong, Shenzhen +4 ยท arXiv
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and re...
#2 TOP PAPER (Score: 72)
Zixiang Wan, Xusheng Yang, Zheng Wang ... ยท Peking University +1 ยท arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
#3 TOP PAPER (Score: 72)
Shibo Wang, Zicheng Zhang, Libo Wang ... ยท Alibaba Group ยท arXiv
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call th...
Tuesday, August 11, 2026
Haoyu Zhang, Zhipeng Li, Xiaoying Tang ... ยท The Chinese University of Hong Kong, Shenzhen +4 ยท arXiv
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and re...
Monday, August 10, 2026
Yan Rong, Fengji Ma, Xu Li ... ยท Kling Team ยท arXiv
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle to achieve these two goals and mainly rely on supervised fine-tuning, yielding sub-optimal performa...
Shuyu Li, Kejun Zhang, Jiahe Lei ... ยท Zhejiang University +3 ยท arXiv
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduc...
Zhanhong He, Hanyu Meng, David Defeng Huang ... ยท University of Western Australia +1 ยท ISMIR 2026
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper s...
Yanqiu Li, Yang Xiao, Jisheng Bai ... ยท arXiv
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet exist...
Saturday, August 08, 2026
Zixiang Wan, Xusheng Yang, Zheng Wang ... ยท Peking University +1 ยท arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J ... ยท Indian Institute of Science (IISc) +2 ยท arXiv
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and...
Friday, August 07, 2026
Shibo Wang, Zicheng Zhang, Libo Wang ... ยท Alibaba Group ยท arXiv
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call th...
Hanke Xie, Haopeng Lin, Jiale Qian, โ˜… Lei Xie ... ยท Shanghai Jiao Tong University +1 ยท arXiv
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit tok...
Thursday, August 06, 2026
Jakub Poฤ‡wiardowski, Mateusz Modrzejewski ยท Warsaw University of Technology ยท arXiv
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extende...
Wednesday, August 05, 2026
Iftach Shoham, Tali Dror, Oren Gal ... ยท Ben-Gurion University of the Negev +1 ยท arXiv
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcr...