Audio ML Papers

Last 7 Days (July 22 - July 29, 2026)

Subcategories: All (20) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (0) | Midi Generation (1) | Generative Conditioning (0) | Other (19)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 75)
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... · Universitat Politècnica de València +2 · arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
#2 TOP PAPER (Score: 74)
Siqian Tong, Xuan Li, Chaozhuo Li ... · Institute of Computing Technology, Chinese Academy of Sciences +8 · arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
#3 TOP PAPER (Score: 74)
Jun Zhan, Chen Yang, Yitian Gong ... · Fudan University +3 · arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Tuesday, July 28, 2026
David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos ... · Universitat Politècnica de València +2 · arXiv
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are of...
Yujian Ma, Jinqiu Sang, Ruizhe Li ... · arXiv
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affect...
Gyeongmin Kim · Hanyang University · arXiv
This is an implementation and measurement study of what it costs to run a streaming speech enhancer on a CPU. We port FastEnhancer-Medium at 48 kHz to faster-enhancer.c, a C runtime with six int8 GEMM tiers selected at initialization, leaving architecture and weights untouched. O...
Stephen Bauer, Sheila Seidel, Shanza Iftikhar ... · Analog Devices, Inc. +4 · INTERSPEECH 2026
Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent lay...
Gyeongmin Kim · Hanyang University · arXiv
Some text-to-speech systems ship a synthesis model and preset style vectors but not the reference encoder that turns audio into such a vector. The model still accepts a style vector; a user with a voice of their own cannot produce one. We solve for that input directly, inverting ...
Monday, July 27, 2026
Mohan Li, Rama Doddipatla, Philip C. Woodland · University of Cambridge +1 · arXiv
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address...
Sunday, July 26, 2026
Jun Zhan, Chen Yang, Yitian Gong ... · Fudan University +3 · arXiv
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to the...
Friday, July 24, 2026
Ziyu Wang, Kun Fang, ★ Yann LeCun · McGill University +4 · arXiv (Submitted to ISMIR 2025 based on template reference)
Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remain...
Yining Yang, Ruogu Chen, Jie Han · University of Alberta +1 · International Society for Music Information Retrieval (ISMIR)
Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction ...
Pengfei Zhang, Biao Tian, Tianxin Xie ... · arXiv
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text le...
Thursday, July 23, 2026
Daniyal Kabir Dar, Arun Ross · Michigan State University · IEEE/IAPR International Joint Conference on Biometrics (IJCB) 2026
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different d...
Songchen Xu, Ting Song, Shaohan Huang ... · Microsoft Research +3 · arXiv
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization ...
Shengkui Zhao, Zexu Pan, Haoxu Wang ... · Alibaba Group +1 · arXiv (preprint)
Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range ...
Hao Zhang, Yiwen Zhao, Yixuan Zhang ... · Tencent Hunyuan +1 · arXiv (preprint)
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based ag...
Viola Negroni, Xin Wang, Wanying Ge, ★ Junichi Yamagishi ... · Politecnico di Milano +1 · arXiv
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is ...
Wednesday, July 22, 2026
Siqian Tong, Xuan Li, Chaozhuo Li ... · Institute of Computing Technology, Chinese Academy of Sciences +8 · arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
Rongshen He, Xinyu Liang, Dekun Chen ... · The Chinese University of Hong Kong, Shenzhen +2 · arXiv
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enab...
Tieyao Zhang, Yuke Liu, Jiaxing Yu ... · Zhejiang University +2 · arXiv (preprint)
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning...
Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak ... · Yandex · Interspeech 2026
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average...
Kaicheng Luo, Xuefei Gong, Yutao Sun ... · Honor Device Co., Ltd. +3 · ASRU 2025
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignmen...