Audio ML Papers

Last 7 Days (July 29 - August 05, 2026)

Subcategories: All (19) | Speech Synthesis (1) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (18)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 75)
Qingjian Lin, Yuxin Li, Haoyang Zhang ... ¡ StepFun +5 ¡ arXiv
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundam...
#2 TOP PAPER (Score: 75)
Yu Zhang, Ruiqi Li, Changhao Pan ... ¡ ByteDance +1 ¡ arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
#3 TOP PAPER (Score: 74)
Fangxu Yu, Tao Feng, Dehai Min ... ¡ Microsoft Research +5 ¡ arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
Tuesday, August 04, 2026
Guanrou Yang, Tian Tan, Qian Chen ... ¡ Shanghai Innovation Institute +1 ¡ arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Jingwei Zhao, Gus Xia, Ziyu Wang ... ¡ Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) +2 ¡ ISMIR 2026
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio ...
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang ... ¡ Tencent AI Lab +2 ¡ arXiv
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict...
Xiang Lin, Tian-Hao Zhang, Chunfeng Wang ... ¡ StepFun (inferred from "Step-Audio-2-mini" and author list context, though explicitly listed as Step in affiliations) +1 ¡ arXiv
Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is challenging because communicative intent may be stated explicitly in lexical content or conveyed more i...
Ye-Xin Lu, Xin Wang, Yang Ai, ★ Zhen-Hua Ling, ★ Junichi Yamagishi ... · University of Science and Technology of China +1 · IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) [Inferred from IEEEtran class and "Journal of Class Files"]
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability....
Monday, August 03, 2026
Yu Zhang, Ruiqi Li, Changhao Pan ... ¡ ByteDance +1 ¡ arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
Fangxu Yu, Tao Feng, Dehai Min ... ¡ Microsoft Research +5 ¡ arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
Zhenghui Guo, Yilin Yang, Yuanbin Man ... ¡ University of Houston +1 ¡ arXiv
Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retained capacity each modality rec...
Sunday, August 02, 2026
Yinhao Bai, Jinming Chen, Yafeng Chen, ★ Yuxuan Wang ... · JD.com · arXiv
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech...
Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen, ★ Yi-Hsuan Yang ... · National Taiwan University +5 · 27th International Society for Music Information Retrieval Conference (ISMIR)
Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alterna...
Saturday, August 01, 2026
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, ★ Jonathan Le Roux ... · Mitsubishi Electric Research Laboratories (MERL) +3 · arXiv (DCASE 2026 Challenge Task 2)
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farth...
Junchuan Zhao, Minh Duc Vu, Bowen Zhang ... ¡ National University of Singapore +1 ¡ arXiv
Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversion problem, potentially requiring cutoff-specific models or retraining when the input bandwidth chang...
Friday, July 31, 2026
Qingjian Lin, Yuxin Li, Haoyang Zhang ... ¡ StepFun +5 ¡ arXiv
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundam...
Yi Luo, Rongzhi Gu, Jixun Yao ¡ ByteDance Seed ¡ arXiv
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with higher frame rates or greater capacity can preserve more signal detail, but they also make streaming gene...
Thursday, July 30, 2026
Xilin Jiang, Riki Shimizu, Sukru Samet Dindar ... ¡ Columbia University +1 ¡ arXiv
Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant ...
Wednesday, July 29, 2026
Jiachen Qian, Junyu Li ¡ City University of Hong Kong +1 ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26)
Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or br...
Junyu Dai, Xiaoyue Duan, Xinyue Fan ... ¡ Alibaba Token Foundry ¡ arXiv
Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion...
Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, ★ Yi-Hsuan Yang ... · Taiwan AI Labs +2 · 27th International Society for Music Information Retrieval (ISMIR)
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit l...
Ke Zhang, Xiaoyang Yu, Haoyu Li ... ¡ The Chinese University of Hong Kong, Shenzhen +4 ¡ Interspeech 2026
The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSe...