Audio ML Papers

Last 7 Days (September 15 - September 22, 2026)

Subcategories: All (14) | Speech Synthesis (1) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (1) | Enhancement (0) | Asr (2) | Llm Audio (8) | Midi Generation (1) | Generative Conditioning (0) | Other (1)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 80)
Lujia Bao, Qian Chen, Luyao Cheng ... · Alibaba Group +1 · arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
#2 TOP PAPER (Score: 76)
Haolin He, Yunfei Chu, Qi Chen ... · Alibaba Group +4 · arXiv
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, with...
#3 TOP PAPER (Score: 75)
Xiuwen Zheng · University of Illinois Urbana-Champaign · arXiv
Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $τ$ that boun...
Monday, September 21, 2026
Lujia Bao, Qian Chen, Luyao Cheng ... · Alibaba Group +1 · arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang, ★ Hung-yi Lee ... · National Taiwan University +1 · arXiv
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can fait...
Sunday, September 20, 2026
Jiaheng Dong, Xiaofeng Yu, Jean Honorio ... · University of Maryland, College Park +4 · AAAI 2027 (Submission)
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTR...
Haoyue Liu, Ye Chen, Zhichao Wang ... · Anonymous (Double-Blind Review) +1 · ICLR 2027 (inferred from template `iclr2027_conference`)
Constraint-following music generation asks a score to satisfy several user-specified properties at once, each checkable programmatically (key, meter, length, range, final note, rhythm, motion and form), yet no existing benchmark isolates this capability. We construct MusicConstra...
Saturday, September 19, 2026
Nishit Anand, Jiaqi Su, Ke Chen, ★ Rithesh Kumar ... · University of Maryland +2 · Interspeech 2026
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralingu...
Jing Peng, Zichao Nie, Zhisheng Zhang ... · Tsinghua University · INTERSPEECH 2026
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To add...
Friday, September 18, 2026
Haolin He, Yunfei Chu, Qi Chen ... · Alibaba Group +4 · arXiv
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, with...
Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang ... · Meta Reality Labs · Interspeech 2026
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearab...
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski · Samsung R&D Institute Poland +1 · Interspeech 2026
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-dev...
Thursday, September 17, 2026
Thomas J Stoll, Ross K Maddox · University of Michigan +2 · arXiv
Computational models of auditory physiology commonly target specific responses or stages of the auditory pathway, limiting their ability to integrate findings across experimental paradigms and neural timescales. We present a foundation model of human auditory electrophysiology: a...
Zifan Guan, Longyu Lu, Junan Zhang ... · The Chinese University of Hong Kong +2 · arXiv
Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspe...
Wednesday, September 16, 2026
Xiuwen Zheng · University of Illinois Urbana-Champaign · arXiv
Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $τ$ that boun...
Mohan Shi, Zilai Wang, Natarajan Balaji Shankar ... · University of California Los Angeles · IEEE SLT 2026
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting t...
Tuesday, September 15, 2026
Rongxiang Wang, Berkin Durmus, Aysegul Orhon ... · Argmax, Inc. +4 · arXiv
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and spea...