Audio ML Papers

Last 7 Days (September 18 - September 25, 2026)

Subcategories: All (10) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (8) | Midi Generation (1) | Generative Conditioning (0) | Other (1)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 80)
Lujia Bao, Qian Chen, Luyao Cheng ... · Alibaba Group +1 · arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
#2 TOP PAPER (Score: 76)
Haolin He, Yunfei Chu, Qi Chen ... · Alibaba Group +4 · arXiv
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, with...
#3 TOP PAPER (Score: 75)
Nishit Anand, Jiaqi Su, Ke Chen, ★ Rithesh Kumar ... · University of Maryland +2 · Interspeech 2026
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralingu...
Thursday, September 24, 2026
Alessandro Bondielli, Lucia Passaro, Serena Auriemma ... · University of Pisa · EMNLP 2026
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic informati...
Monday, September 21, 2026
Lujia Bao, Qian Chen, Luyao Cheng ... · Alibaba Group +1 · arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang, ★ Hung-yi Lee ... · National Taiwan University +1 · arXiv
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can fait...
Sunday, September 20, 2026
Jiaheng Dong, Xiaofeng Yu, Jean Honorio ... · University of Maryland, College Park +4 · AAAI 2027 (Submission)
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTR...
Haoyue Liu, Ye Chen, Zhichao Wang ... · Anonymous (Double-Blind Review) +1 · ICLR 2027 (inferred from template `iclr2027_conference`)
Constraint-following music generation asks a score to satisfy several user-specified properties at once, each checkable programmatically (key, meter, length, range, final note, rhythm, motion and form), yet no existing benchmark isolates this capability. We construct MusicConstra...
Saturday, September 19, 2026
Nishit Anand, Jiaqi Su, Ke Chen, ★ Rithesh Kumar ... · University of Maryland +2 · Interspeech 2026
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralingu...
Jing Peng, Zichao Nie, Zhisheng Zhang ... · Tsinghua University · INTERSPEECH 2026
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To add...
Friday, September 18, 2026
Haolin He, Yunfei Chu, Qi Chen ... · Alibaba Group +4 · arXiv
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, with...
Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang ... · Meta Reality Labs · Interspeech 2026
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearab...
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski · Samsung R&D Institute Poland +1 · Interspeech 2026
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-dev...