Audio ML Papers

Last 7 Days (August 04 - August 11, 2026)

Subcategories: All (15) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (14)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 72)
Guanrou Yang, Tian Tan, Qian Chen ... ¡ Shanghai Innovation Institute +1 ¡ arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
#2 TOP PAPER (Score: 72)
Jingwei Zhao, Gus Xia, Ziyu Wang ... ¡ Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) +2 ¡ ISMIR 2026
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio ...
#3 TOP PAPER (Score: 72)
Zixiang Wan, Xusheng Yang, Zheng Wang ... ¡ Peking University +1 ¡ arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
Monday, August 10, 2026
Yan Rong, Fengji Ma, Xu Li ... ¡ Kling Team ¡ arXiv
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle to achieve these two goals and mainly rely on supervised fine-tuning, yielding sub-optimal performa...
Shuyu Li, Kejun Zhang, Jiahe Lei ... ¡ Zhejiang University +3 ¡ arXiv
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduc...
Zhanhong He, Hanyu Meng, David Defeng Huang ... ¡ University of Western Australia +1 ¡ ISMIR 2026
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper s...
Yanqiu Li, Yang Xiao, Jisheng Bai ... ¡ arXiv
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet exist...
Saturday, August 08, 2026
Zixiang Wan, Xusheng Yang, Zheng Wang ... ¡ Peking University +1 ¡ arXiv
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows tha...
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J ... ¡ Indian Institute of Science (IISc) +2 ¡ arXiv
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and...
Friday, August 07, 2026
Shibo Wang, Zicheng Zhang, Libo Wang ... ¡ Alibaba Group ¡ arXiv
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call th...
Hanke Xie, Haopeng Lin, Jiale Qian, ★ Lei Xie ... · Shanghai Jiao Tong University +1 · arXiv
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit tok...
Thursday, August 06, 2026
Jakub Poćwiardowski, Mateusz Modrzejewski · Warsaw University of Technology · arXiv
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extende...
Wednesday, August 05, 2026
Iftach Shoham, Tali Dror, Oren Gal ... ¡ Ben-Gurion University of the Negev +1 ¡ arXiv
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcr...
Tuesday, August 04, 2026
Guanrou Yang, Tian Tan, Qian Chen ... ¡ Shanghai Innovation Institute +1 ¡ arXiv
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead...
Jingwei Zhao, Gus Xia, Ziyu Wang ... ¡ Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) +2 ¡ ISMIR 2026
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio ...
Zixun Guo, ★ Simon Dixon · Queen Mary University of London +1 · ISMIR 2026 · ISMIR 2026
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrel...
Junhao Chen, Mingjin Chen, Jingjia Mao ... ¡ Unknown (Affiliations not explicitly named in text, only numbered) +1 ¡ arXiv
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only t...
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang ... ¡ Tencent AI Lab +2 ¡ arXiv
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict...