Audio ML Papers

Last 7 Days (August 29 - September 05, 2026)

Subcategories: All (14) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (3) | Quality Evaluation (0) | Enhancement (1) | Asr (2) | Llm Audio (2) | Midi Generation (0) | Generative Conditioning (0) | Other (6)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 81)
Dongwook Lee, Sangkwon Park, Eunwoo Song ... ยท Seoul National University +2 ยท EMNLP 2026
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the f...
#2 TOP PAPER (Score: 80)
Yujie Tu, Zhiliang Peng, Jianwei Yu ... ยท Microsoft Research +2 ยท arXiv
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, mak...
#3 TOP PAPER (Score: 76)
Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones ยท University of Oxford ยท arXiv
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common m...
Friday, September 04, 2026
Yusuke Oumi, Yuto Shibata, Go Irie ... ยท Keio University +2 ยท ECCV 2026
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superpos...
Ke Lei, Chenyuhao Wen, Yu Zhang ... ยท Zhejiang University +1 ยท arXiv
Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in fir...
Thursday, September 03, 2026
Linyi Jiang, Silvery D. Fu, Yifei Zhu ยท Shanghai Jiao Tong University +1 ยท ACM SOSP 2026
Voice AI applications are gaining popularity as advances in large language models (LLMs) enable more natural and accessible spoken interactions. Serving these applications requires accounting not only for what users say, but also for how they speak (e.g., speaking rate) and the c...
Sanyuan Chen, Min-Jae Hwang, Sho Inoue, โ˜… Juan Pino, โ˜… Wei-Ning Hsu ... ยท arXiv
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. F...
Wednesday, September 02, 2026
Yujie Tu, Zhiliang Peng, Jianwei Yu ... ยท Microsoft Research +2 ยท arXiv
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, mak...
Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones ยท University of Oxford ยท arXiv
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common m...
Mojtaba Nafez, Mobina Poulaei, Kiarash Kiani Feriz ... ยท Sharif University of Technology +2 ยท arXiv
Automatic Speech Recognition (ASR) systems are widely deployed in safety-critical settings but remain vulnerable to data-poisoning backdoor attacks. Existing ASR backdoors typically use phrase-level triggers paired with a fixed target sentence, creating strong artifacts (e.g., re...
Monday, August 31, 2026
Zixun Guo, Calvin Murdock, Sanjeel Parekh, โ˜… Simon Dixon ... ยท Meta +1 ยท EMNLP 2026 ยท EMNLP 2026
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-trai...
Chanhee Cho, Junhyuk Choi, Bugeun Kim ยท Chung-Ang University ยท EMNLP 2026
Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operat...
Yunjie Zhou, Yuheng Huang, Diqun Yan ยท Ningbo University +1 ยท INTERSPEECH 2026
Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incom...
Randy Frans Fela, Pejman Mowlaee ยท GN Group ยท EMNLP 2026
Speech enhancement (SE) is commonly applied as a preprocessing step in spoken AI pipelines under the assumption that better audio quality improves downstream task performance. Whether SE-induced distortions propagate to downstream LLM task performance remains an open question. We...
Sunday, August 30, 2026
Lingfeng Yao, Chenpei Huang, Xingke Yang ... ยท University of Houston +2 ยท EMNLP 2026
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acous...
Saturday, August 29, 2026
Dongwook Lee, Sangkwon Park, Eunwoo Song ... ยท Seoul National University +2 ยท EMNLP 2026
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the f...
Zeyang Song, Tianchi Liu, Tianrui Wang ... ยท National University of Singapore +7 ยท EMNLP 2026
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filt...