Audio ML Papers

Last 7 Days (July 24 - July 31, 2026)

Subcategories: All (17) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (17)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 79)
Bajian Xiang, Cheng Wen, Han Zhao ... ¡ Alibaba Group ¡ arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
#2 TOP PAPER (Score: 75)
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... ¡ Universitat Politècnica de València +2 ¡ arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
#3 TOP PAPER (Score: 74)
Jun Zhan, Chen Yang, Yitian Gong ... ¡ Fudan University +3 ¡ arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Thursday, July 30, 2026
Xilin Jiang, Riki Shimizu, Sukru Samet Dindar ... ¡ Columbia University +1 ¡ arXiv
Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant ...
Wednesday, July 29, 2026
Jiachen Qian, Junyu Li ¡ City University of Hong Kong +1 ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26)
Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or br...
Junyu Dai, Xiaoyue Duan, Xinyue Fan ... ¡ Alibaba Token Foundry ¡ arXiv
Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion...
Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, ★ Yi-Hsuan Yang ... · Taiwan AI Labs +2 · 27th International Society for Music Information Retrieval (ISMIR)
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit l...
Ke Zhang, Xiaoyang Yu, Haoyu Li ... ¡ The Chinese University of Hong Kong, Shenzhen +4 ¡ Interspeech 2026
The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSe...
Tuesday, July 28, 2026
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... ¡ Universitat Politècnica de València +2 ¡ arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
Yujian Ma, Jinqiu Sang, Ruizhe Li ... ¡ arXiv
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affect...
Gyeongmin Kim ¡ Hanyang University ¡ arXiv
This is an implementation and measurement study of what it costs to run a streaming speech enhancer on a CPU. We port FastEnhancer-Medium at 48 kHz to faster-enhancer.c, a C runtime with six int8 GEMM tiers selected at initialization, leaving architecture and weights untouched. O...
Stephen Bauer, Sheila Seidel, Shanza Iftikhar ... ¡ Analog Devices, Inc. +4 ¡ INTERSPEECH 2026
Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent lay...
Gyeongmin Kim ¡ Hanyang University ¡ arXiv
Some text-to-speech systems ship a synthesis model and preset style vectors but not the reference encoder that turns audio into such a vector. The model still accepts a style vector; a user with a voice of their own cannot produce one. We solve for that input directly, inverting ...
Monday, July 27, 2026
Bajian Xiang, Cheng Wen, Han Zhao ... ¡ Alibaba Group ¡ arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
Mohan Li, Rama Doddipatla, Philip C. Woodland ¡ University of Cambridge +1 ¡ arXiv
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address...
Sunday, July 26, 2026
Dongseong Hwang, Prasanth Yadla, Kaan Elgin, ★ Chung-Cheng Chiu ... · Apple Inc. · arXiv
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detoke...
Jun Zhan, Chen Yang, Yitian Gong ... ¡ Fudan University +3 ¡ arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Friday, July 24, 2026
Ziyu Wang, Kun Fang, ★ Yann LeCun · McGill University +4 · arXiv (Submitted to ISMIR 2025 based on template reference)
Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remain...
Yining Yang, Ruogu Chen, Jie Han ¡ University of Alberta +1 ¡ International Society for Music Information Retrieval (ISMIR)
Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction ...
Pengfei Zhang, Biao Tian, Tianxin Xie ... ¡ arXiv
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text le...