Audio ML Papers

Last 7 Days (August 19 - August 26, 2026)

Subcategories: All (13) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (1) | Asr (2) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (10)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 72)
Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer ... · ETH Zurich · arXiv
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every syste...
#2 TOP PAPER (Score: 72)
Yi Yuan, ★ Xubo Liu, ★ Haohe Liu, ★ Mark D. Plumbley, ★ Wenwu Wang ... · University of Surrey +2 · arXiv (Preprint)
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-...
#3 TOP PAPER (Score: 71)
Yinming Huang, Shuyuan Tu, Xi Yan ... · Fudan University +3 · arXiv
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these me...
Tuesday, August 25, 2026
Fulin Wu, Zhong-Qiu Wang · unknown (Affiliations not explicitly stated in text, authors are Fulin Wu and Zhong-Qiu Wang) +1 · arXiv
We propose $\textit{recursive encoder and decoder}$ (RED) for building a single deep neural network (DNN) model that can separate multi-speaker mixtures containing unknown numbers of speakers and variable numbers of microphones arranged in an unknown geometry, a task that has not...
Haoyang Li, Chenglin Xu, Junchuan Zhao ... · Nanyang Technological University +3 · arXiv (Preprint)
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on pre...
Qingyu Luo, Peng Zhang, ★ Wenwu Wang ... · University of Surrey · INTERSPEECH 2026
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an ...
Junjie Li, Xuelong Geng, Kun Xie, ★ Lei Xie ... · Xiaohongshu · arXiv
⚠ Prompt injection detected in this paper

The text contains content addressed to automated reviewers. Its score has been penalised, and the ranking below should not be trusted.

[system_override] "ruct TTS and speech editing require the shared LLM to learn new instruction-conditioned output patterns, we retain a peak learning rate"
A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long-context modeling, whereas speech ge...
Monday, August 24, 2026
Songchun Zhang, Yaowei Li, Junhao Zhuang ... · JD (Joy Future Academy) +4 · arXiv
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observe...
Tianchi Liu, Zeyang Song, Tianrui Wang ... · LIGHTSPEED +3 · EMNLP 2026 (Main Conference)
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding p...
Sunday, August 23, 2026
Yize Li, Ningyuan Yang, Sile Yin ... · Northeastern University +2 · arXiv
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event...
Saturday, August 22, 2026
Yi Yuan, ★ Xubo Liu, ★ Haohe Liu, ★ Mark D. Plumbley, ★ Wenwu Wang ... · University of Surrey +2 · arXiv (Preprint)
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-...
Friday, August 21, 2026
Vladimir Bataev, Lilit Grigoryan, Andrei Andrusenko ... · NVIDIA · arXiv
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, they often do not address the ...
Thursday, August 20, 2026
Theo Lebryk, David Ayllon, Alice Baird ... · arXiv
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for...
Umberto Cappellazzo, ★ Xubo Liu, Stavros Petridis ... · Imperial College London +1 · arXiv
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance. A markedly different pre-training philosophy underpins the most in...
Wednesday, August 19, 2026
Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer ... · ETH Zurich · arXiv
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every syste...
Yinming Huang, Shuyuan Tu, Xi Yan ... · Fudan University +3 · arXiv
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these me...