Audio ML Papers

Last 7 Days (August 01 - August 08, 2026)

Subcategories: All (15) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (15)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 75)
Yu Zhang, Ruiqi Li, Changhao Pan ... · ByteDance +1 · arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
#2 TOP PAPER (Score: 74)
Fangxu Yu, Tao Feng, Dehai Min ... · Microsoft Research +5 · arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
#3 TOP PAPER (Score: 72)
Guanrou Yang, Tian Tan, Qian Chen ... · Shanghai Innovation Institute +1 · arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Friday, August 07, 2026
Hanke Xie, Haopeng Lin, Jiale Qian, ★ Lei Xie ... · Shanghai Jiao Tong University +1 · arXiv
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit tok...
Thursday, August 06, 2026
Jakub Poćwiardowski, Mateusz Modrzejewski · Warsaw University of Technology · arXiv
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extende...
Wednesday, August 05, 2026
Iftach Shoham, Tali Dror, Oren Gal ... · Ben-Gurion University of the Negev +1 · arXiv
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcr...
Tuesday, August 04, 2026
Guanrou Yang, Tian Tan, Qian Chen ... · Shanghai Innovation Institute +1 · arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Jingwei Zhao, Gus Xia, Ziyu Wang ... · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) +2 · ISMIR 2026
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio ...
Zixun Guo, ★ Simon Dixon · Queen Mary University of London +1 · ISMIR 2026 · ISMIR 2026
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrel...
Junhao Chen, Mingjin Chen, Jingjia Mao ... · Unknown (Affiliations not explicitly named in text, only numbered) +1 · arXiv
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only t...
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang ... · Tencent AI Lab +2 · arXiv
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict...
Monday, August 03, 2026
Yu Zhang, Ruiqi Li, Changhao Pan ... · ByteDance +1 · arXiv
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, supp...
Fangxu Yu, Tao Feng, Dehai Min ... · Microsoft Research +5 · arXiv
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and le...
Zhenghui Guo, Yilin Yang, Yuanbin Man ... · University of Houston +1 · arXiv
Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retained capacity each modality rec...
Sunday, August 02, 2026
Yinhao Bai, Jinming Chen, Yafeng Chen, ★ Yuxuan Wang ... · JD.com · arXiv
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech...
Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen, ★ Yi-Hsuan Yang ... · National Taiwan University +5 · 27th International Society for Music Information Retrieval Conference (ISMIR)
Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alterna...
Saturday, August 01, 2026
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, ★ Jonathan Le Roux ... · Mitsubishi Electric Research Laboratories (MERL) +3 · arXiv (DCASE 2026 Challenge Task 2)
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farth...
Junchuan Zhao, Minh Duc Vu, Bowen Zhang ... · National University of Singapore +1 · arXiv
Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversion problem, potentially requiring cutoff-specific models or retraining when the input bandwidth chang...