Audio ML Papers

Last 7 Days (July 26 - August 02, 2026)

Subcategories: All (18) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (18)
← Previous Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 79)
Bajian Xiang, Cheng Wen, Han Zhao ... · Alibaba Group · arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
#2 TOP PAPER (Score: 75)
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... · Universitat Politècnica de València +2 · arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
#3 TOP PAPER (Score: 75)
Qingjian Lin, Yuxin Li, Haoyang Zhang ... · StepFun +5 · arXiv
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundam...
Saturday, August 01, 2026
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, ★ Jonathan Le Roux ... · Mitsubishi Electric Research Laboratories (MERL) +3 · arXiv (DCASE 2026 Challenge Task 2)
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farth...
Junchuan Zhao, Minh Duc Vu, Bowen Zhang ... · National University of Singapore +1 · arXiv
Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversion problem, potentially requiring cutoff-specific models or retraining when the input bandwidth chang...
Friday, July 31, 2026
Qingjian Lin, Yuxin Li, Haoyang Zhang ... · StepFun +5 · arXiv
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundam...
Yi Luo, Rongzhi Gu, Jixun Yao · ByteDance Seed · arXiv
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with higher frame rates or greater capacity can preserve more signal detail, but they also make streaming gene...
Thursday, July 30, 2026
Xilin Jiang, Riki Shimizu, Sukru Samet Dindar ... · Columbia University +1 · arXiv
Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant ...
Wednesday, July 29, 2026
Jiachen Qian, Junyu Li · City University of Hong Kong +1 · Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil · Proceedings of the 34th ACM International Conference on Multimedia (MM '26)
Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or br...
Junyu Dai, Xiaoyue Duan, Xinyue Fan ... · Alibaba Token Foundry · arXiv
Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion...
Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, ★ Yi-Hsuan Yang ... · Taiwan AI Labs +2 · 27th International Society for Music Information Retrieval (ISMIR)
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit l...
Ke Zhang, Xiaoyang Yu, Haoyu Li ... · The Chinese University of Hong Kong, Shenzhen +4 · Interspeech 2026
The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSe...
Tuesday, July 28, 2026
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... · Universitat Politècnica de València +2 · arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
Yujian Ma, Jinqiu Sang, Ruizhe Li ... · arXiv
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affect...
Gyeongmin Kim · Hanyang University · arXiv
This is an implementation and measurement study of what it costs to run a streaming speech enhancer on a CPU. We port FastEnhancer-Medium at 48 kHz to faster-enhancer.c, a C runtime with six int8 GEMM tiers selected at initialization, leaving architecture and weights untouched. O...
Stephen Bauer, Sheila Seidel, Shanza Iftikhar ... · Analog Devices, Inc. +4 · INTERSPEECH 2026
Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent lay...
Gyeongmin Kim · Hanyang University · arXiv
Some text-to-speech systems ship a synthesis model and preset style vectors but not the reference encoder that turns audio into such a vector. The model still accepts a style vector; a user with a voice of their own cannot produce one. We solve for that input directly, inverting ...
Monday, July 27, 2026
Bajian Xiang, Cheng Wen, Han Zhao ... · Alibaba Group · arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
Mohan Li, Rama Doddipatla, Philip C. Woodland · University of Cambridge +1 · arXiv
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address...
Sunday, July 26, 2026
Dongseong Hwang, Prasanth Yadla, Kaan Elgin, ★ Chung-Cheng Chiu ... · Apple Inc. · arXiv
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detoke...
Jun Zhan, Chen Yang, Yitian Gong ... · Fudan University +3 · arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...