Audio ML Papers

Last 7 Days (September 03 - September 10, 2026)

Subcategories: All (11) | Speech Synthesis (2) | Music Synthesis (0) | Ambient Synthesis (1) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (3) | Midi Generation (0) | Generative Conditioning (0) | Other (4)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 81)
Ziyang Ma, Zhikang Niu, Wenming Tu ... · Shanghai Jiao Tong University · arXiv
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances...
#2 TOP PAPER (Score: 79)
Song-ha Jo, Sehyun Lee, Soyoon Kim ... · Seoul National University +2 · arXiv
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whi...
#3 TOP PAPER (Score: 75)
Bella Godiva, Yeonju Kim, Yong Man Ro · KAIST · EMNLP 2026
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues ...
Wednesday, September 09, 2026
Yupei Li, Qiyang Sun, Mohamed Mady ... · Imperial College London +4 · AACL 2026
Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than g...
Mingyu Zhao, Zhiyong Wu · Tsinghua University +1 · NCMMSC 2026
We present UniStream, a fully causal 48 kHz neural audio codec for streaming speech, music, and environmental sounds. At its core is Multi-Expert Residual Vector Quantization (ME-RVQ), which replaces the single shared codebook in each residual quantization layer with four expert ...
Tuesday, September 08, 2026
Ziyang Ma, Zhikang Niu, Wenming Tu ... · Shanghai Jiao Tong University · arXiv
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances...
Bella Godiva, Yeonju Kim, Yong Man Ro · KAIST · EMNLP 2026
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues ...
Orantqing, Shengpeng Ji, Junlong Tong ... · Tencent +4 · arXiv
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities,...
Saturday, September 05, 2026
Song-ha Jo, Sehyun Lee, Soyoon Kim ... · Seoul National University +2 · arXiv
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whi...
Friday, September 04, 2026
Yusuke Oumi, Yuto Shibata, Go Irie ... · Keio University +2 · ECCV 2026
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superpos...
Ke Lei, Chenyuhao Wen, Yu Zhang ... · Zhejiang University +1 · arXiv
Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in fir...
Thursday, September 03, 2026
Linyi Jiang, Silvery D. Fu, Yifei Zhu · Shanghai Jiao Tong University +1 · ACM SOSP 2026
Voice AI applications are gaining popularity as advances in large language models (LLMs) enable more natural and accessible spoken interactions. Serving these applications requires accounting not only for what users say, but also for how they speak (e.g., speaking rate) and the c...
Sagi Polaczek, Noa Kraicer, Gal Metzer ... · Tel Aviv University +1 · arXiv
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three ...
Sanyuan Chen, Min-Jae Hwang, Sho Inoue, ★ Juan Pino, ★ Wei-Ning Hsu ... · arXiv
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. F...