Audio ML Papers

Last 7 Days (August 17 - August 24, 2026)

Subcategories: All (13) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (1) | Asr (2) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (10)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 72)
Feiyu Shen, Kun Xie, Yichen Wu, ★ Lei Xie ... · Xiaohongshu · arXiv
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlle...
#2 TOP PAPER (Score: 72)
Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer ... · ETH Zurich · arXiv
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every syste...
#3 TOP PAPER (Score: 72)
Yi Yuan, ★ Xubo Liu, ★ Haohe Liu, ★ Mark D. Plumbley, ★ Wenwu Wang ... · University of Surrey +2 · arXiv (Preprint)
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-...
Sunday, August 23, 2026
Yize Li, Ningyuan Yang, Sile Yin ... · Northeastern University +2 · arXiv
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event...
Saturday, August 22, 2026
Yi Yuan, ★ Xubo Liu, ★ Haohe Liu, ★ Mark D. Plumbley, ★ Wenwu Wang ... · University of Surrey +2 · arXiv (Preprint)
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-...
Friday, August 21, 2026
Vladimir Bataev, Lilit Grigoryan, Andrei Andrusenko ... · NVIDIA · arXiv
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, they often do not address the ...
Thursday, August 20, 2026
Theo Lebryk, David Ayllon, Alice Baird ... · arXiv
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for...
Umberto Cappellazzo, ★ Xubo Liu, Stavros Petridis ... · Imperial College London +1 · arXiv
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance. A markedly different pre-training philosophy underpins the most in...
Wednesday, August 19, 2026
Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer ... · ETH Zurich · arXiv
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every syste...
Yinming Huang, Shuyuan Tu, Xi Yan ... · Fudan University +3 · arXiv
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these me...
Tuesday, August 18, 2026
Feiyu Shen, Kun Xie, Yichen Wu, ★ Lei Xie ... · Xiaohongshu · arXiv
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlle...
Monday, August 17, 2026
Tony Alex, Wish Suharitdamrong, Sara Atito ... · University of Surrey +1 · arXiv
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of se...
Hanlin Zhang, Daxin Tan, Dehua Tao ... · City University of Hong Kong +2 · arXiv
Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training proce...
Fengji Ma, Yan Rong, Xu Li ... · Kling Team · arXiv
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, ...
Chen-An Li, ★ Hung-yi Lee · National Taiwan University +1 · arXiv
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including...
Tao Feng, Xu Li, Xiangyang Luo ... · Kuaishou Technology · arXiv
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily...