Audio ML Papers

Last 7 Days (August 27 - September 03, 2026)

Subcategories: All (8) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (2) | Quality Evaluation (0) | Enhancement (1) | Asr (2) | Llm Audio (1) | Midi Generation (0) | Generative Conditioning (0) | Other (2)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 81)
Dongwook Lee, Sangkwon Park, Eunwoo Song ... ยท Seoul National University +2 ยท EMNLP 2026
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the f...
#2 TOP PAPER (Score: 80)
Yujie Tu, Zhiliang Peng, Jianwei Yu ... ยท Microsoft Research +2 ยท arXiv
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, mak...
#3 TOP PAPER (Score: 76)
Lingfeng Yao, Chenpei Huang, Xingke Yang ... ยท University of Houston +2 ยท EMNLP 2026
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acous...
Wednesday, September 02, 2026
Yujie Tu, Zhiliang Peng, Jianwei Yu ... ยท Microsoft Research +2 ยท arXiv
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, mak...
Monday, August 31, 2026
Zixun Guo, Calvin Murdock, Sanjeel Parekh, โ˜… Simon Dixon ... ยท Meta +1 ยท EMNLP 2026 ยท EMNLP 2026
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-trai...
Chanhee Cho, Junhyuk Choi, Bugeun Kim ยท Chung-Ang University ยท EMNLP 2026
Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operat...
Yunjie Zhou, Yuheng Huang, Diqun Yan ยท Ningbo University +1 ยท INTERSPEECH 2026
Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incom...
Randy Frans Fela, Pejman Mowlaee ยท GN Group ยท EMNLP 2026
Speech enhancement (SE) is commonly applied as a preprocessing step in spoken AI pipelines under the assumption that better audio quality improves downstream task performance. Whether SE-induced distortions propagate to downstream LLM task performance remains an open question. We...
Sunday, August 30, 2026
Lingfeng Yao, Chenpei Huang, Xingke Yang ... ยท University of Houston +2 ยท EMNLP 2026
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acous...
Saturday, August 29, 2026
Dongwook Lee, Sangkwon Park, Eunwoo Song ... ยท Seoul National University +2 ยท EMNLP 2026
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the f...
Zeyang Song, Tianchi Liu, Tianrui Wang ... ยท National University of Singapore +7 ยท EMNLP 2026
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filt...