Audio ML Papers

Week of August 02 - August 09, 2026

Subcategories: All (16) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (15)
← Previous Week | Current Week →

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 75)
Yu Zhang, Ruiqi Li, Changhao Pan ... ยท ByteDance +1 ยท arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
#2 TOP PAPER (Score: 74)
Fangxu Yu, Tao Feng, Dehai Min ... ยท Microsoft Research +5 ยท arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
#3 TOP PAPER (Score: 72)
Guanrou Yang, Tian Tan, Qian Chen ... ยท Shanghai Innovation Institute +1 ยท arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Saturday, August 08, 2026
Zixiang Wan, Xusheng Yang, Zheng Wang ... ยท Peking University +1 ยท arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J ... ยท Indian Institute of Science (IISc) +2 ยท arXiv
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and...
Friday, August 07, 2026
Shibo Wang, Zicheng Zhang, Libo Wang ... ยท Alibaba Group ยท arXiv
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call th...
Hanke Xie, Haopeng Lin, Jiale Qian, โ˜… Lei Xie ... ยท Shanghai Jiao Tong University +1 ยท arXiv
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit tok...
Thursday, August 06, 2026
Jakub Poฤ‡wiardowski, Mateusz Modrzejewski ยท Warsaw University of Technology ยท arXiv
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extende...
Wednesday, August 05, 2026
Iftach Shoham, Tali Dror, Oren Gal ... ยท Ben-Gurion University of the Negev +1 ยท arXiv
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcr...
Tuesday, August 04, 2026
Guanrou Yang, Tian Tan, Qian Chen ... ยท Shanghai Innovation Institute +1 ยท arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Jingwei Zhao, Gus Xia, Ziyu Wang ... ยท Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) +2 ยท ISMIR 2026
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio ...
Zixun Guo, โ˜… Simon Dixon ยท Queen Mary University of London +1 ยท ISMIR 2026 ยท ISMIR 2026
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrel...
Junhao Chen, Mingjin Chen, Jingjia Mao ... ยท Unknown (Affiliations not explicitly named in text, only numbered) +1 ยท arXiv
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only t...
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang ... ยท Tencent AI Lab +2 ยท arXiv
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict...
Monday, August 03, 2026
Yu Zhang, Ruiqi Li, Changhao Pan ... ยท ByteDance +1 ยท arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
Fangxu Yu, Tao Feng, Dehai Min ... ยท Microsoft Research +5 ยท arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
Zhenghui Guo, Yilin Yang, Yuanbin Man ... ยท University of Houston +1 ยท arXiv
Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retained capacity each modality rec...
Sunday, August 02, 2026
Yinhao Bai, Jinming Chen, Yafeng Chen, โ˜… Yuxuan Wang ... ยท JD.com ยท arXiv
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech...
Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen, โ˜… Yi-Hsuan Yang ... ยท National Taiwan University +5 ยท 27th International Society for Music Information Retrieval Conference (ISMIR)
Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alterna...