Audio ML Papers

Last 7 Days (July 23 - July 30, 2026)

Subcategories: All (20) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (20)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 79)
Bajian Xiang, Cheng Wen, Han Zhao ... ¡ Alibaba Group ¡ arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
#2 TOP PAPER (Score: 75)
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... ¡ Universitat Politècnica de València +2 ¡ arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
#3 TOP PAPER (Score: 74)
Jun Zhan, Chen Yang, Yitian Gong ... ¡ Fudan University +3 ¡ arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Wednesday, July 29, 2026
Jiachen Qian, Junyu Li ¡ City University of Hong Kong +1 ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil ¡ Proceedings of the 34th ACM International Conference on Multimedia (MM '26)
Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or br...
Junyu Dai, Xiaoyue Duan, Xinyue Fan ... ¡ Alibaba Token Foundry ¡ arXiv
Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion...
Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, ★ Yi-Hsuan Yang ... · Taiwan AI Labs +2 · 27th International Society for Music Information Retrieval (ISMIR)
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit l...
Tuesday, July 28, 2026
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... ¡ Universitat Politècnica de València +2 ¡ arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
Yujian Ma, Jinqiu Sang, Ruizhe Li ... ¡ arXiv
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affect...
Gyeongmin Kim ¡ Hanyang University ¡ arXiv
This is an implementation and measurement study of what it costs to run a streaming speech enhancer on a CPU. We port FastEnhancer-Medium at 48 kHz to faster-enhancer.c, a C runtime with six int8 GEMM tiers selected at initialization, leaving architecture and weights untouched. O...
Stephen Bauer, Sheila Seidel, Shanza Iftikhar ... ¡ Analog Devices, Inc. +4 ¡ INTERSPEECH 2026
Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent lay...
Gyeongmin Kim ¡ Hanyang University ¡ arXiv
Some text-to-speech systems ship a synthesis model and preset style vectors but not the reference encoder that turns audio into such a vector. The model still accepts a style vector; a user with a voice of their own cannot produce one. We solve for that input directly, inverting ...
Monday, July 27, 2026
Bajian Xiang, Cheng Wen, Han Zhao ... ¡ Alibaba Group ¡ arXiv
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~...
Mohan Li, Rama Doddipatla, Philip C. Woodland ¡ University of Cambridge +1 ¡ arXiv
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address...
Sunday, July 26, 2026
Dongseong Hwang, Prasanth Yadla, Kaan Elgin, ★ Chung-Cheng Chiu ... · Apple Inc. · arXiv
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detoke...
Jun Zhan, Chen Yang, Yitian Gong ... ¡ Fudan University +3 ¡ arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Friday, July 24, 2026
Ziyu Wang, Kun Fang, ★ Yann LeCun · McGill University +4 · arXiv (Submitted to ISMIR 2025 based on template reference)
Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remain...
Yining Yang, Ruogu Chen, Jie Han ¡ University of Alberta +1 ¡ International Society for Music Information Retrieval (ISMIR)
Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction ...
Pengfei Zhang, Biao Tian, Tianxin Xie ... ¡ arXiv
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text le...
Thursday, July 23, 2026
Daniyal Kabir Dar, Arun Ross ¡ Michigan State University ¡ IEEE/IAPR International Joint Conference on Biometrics (IJCB) 2026
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different d...
Songchen Xu, Ting Song, Shaohan Huang ... ¡ Microsoft Research +3 ¡ arXiv
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization ...
Shengkui Zhao, Zexu Pan, Haoxu Wang ... ¡ Alibaba Group +1 ¡ arXiv (preprint)
Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range ...
Hao Zhang, Yiwen Zhao, Yixuan Zhang ... ¡ Tencent Hunyuan +1 ¡ arXiv (preprint)
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based ag...
Viola Negroni, Xin Wang, Wanying Ge, ★ Junichi Yamagishi ... · Politecnico di Milano +1 · arXiv
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is ...