Audio ML Papers

Last 7 Days (September 24 - October 01, 2026)

Subcategories: All (24) | Speech Synthesis (5) | Music Synthesis (0) | Ambient Synthesis (1) | Quality Evaluation (0) | Enhancement (2) | Asr (1) | Llm Audio (7) | Midi Generation (0) | Generative Conditioning (1) | Other (7)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 84)
Ruibin Yuan, Jiahao Pan, Junyan Jiang ... ยท HKGAI ยท arXiv
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through...
#2 TOP PAPER (Score: 83)
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann ยท MIT CSAIL +2 ยท NeurIPS 2026
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound alon...
#3 TOP PAPER (Score: 81)
Yayue Deng, Dingdong Wang, Yuxuan Hu ... ยท The Chinese University of Hong Kong +1 ยท EMNLP 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target s...
Wednesday, September 30, 2026
Kyoungjun Park, Yunzhe Li, Lili Qiu ยท The University of Texas at Austin ยท arXiv
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their rank...
Shengbo Cai, Zhisheng Zhang, Zichao Nie ... ยท Tsinghua University ยท arXiv
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Expe...
Tuesday, September 29, 2026
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann ยท MIT CSAIL +2 ยท NeurIPS 2026
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound alon...
Yunzhe Li, Kyoungjun Park, Hongzi Zhu ... ยท The University of Texas at Austin +1 ยท arXiv
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and veri...
Duowen Chen, Jinjin He, Gouthaman KV ... ยท Georgia Institute of Technology +1 ยท NeurIPS 2026
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representat...
Jisoo Park, Seonghak Lee, Hyojin Park ... ยท Chung-Ang University +1 ยท NeurIPS 2026
Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fideli...
Kuan-Po Huang, โ˜… Haohe Liu, Puyuan Peng, โ˜… Hung-yi Lee ... ยท Meta AI (Reality Labs / FAIR) +3 ยท arXiv
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free app...
Monday, September 28, 2026
Wenyi Yu, Siyin Wang, Terumi Chiba ... ยท Tsinghua University +1 ยท arXiv (Preprint)
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirement...
Tianxin Xie, Pengfei Zhang, Kai Jiang ... ยท The Hong Kong University of Science and Technology (Guangzhou) ยท arXiv
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for trai...
Donghang Wu, Haoyang Zhang, Yizhou Peng ... ยท Nanyang Technological University +2 ยท arXiv
Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like halluc...
Sunday, September 27, 2026
Ruibin Yuan, Jiahao Pan, Junyan Jiang ... ยท HKGAI ยท arXiv
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through...
Yayue Deng, Dingdong Wang, Yuxuan Hu ... ยท The Chinese University of Hong Kong +1 ยท EMNLP 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target s...
Yifan Yang, Xiaoyu Yang, Zengrui Jin, โ˜… Yuxuan Wang ... ยท Alibaba Group +6 ยท arXiv
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohi...
Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu ยท University of Zurich and ETH Zurich +4 ยท arXiv
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation...
Saturday, September 26, 2026
Francesco Brigante, Luca Cerovaz, Davide Marincione ... ยท Sapienza University of Rome +3 ยท arXiv
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference spee...
Ambuj Mehrish, Abhinaba Roy, Alex Ivanov ... ยท Ca' Foscari University of Venice +3 ยท arXiv
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identi...
Yang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden ... ยท arXiv
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be r...
Friday, September 25, 2026
Jihoo Jung, Youngjoon Jang, โ˜… Joon Son Chung ยท KAIST +2 ยท NeurIPS 2026
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate ho...
Yotaro Kubo, Qi Sun, Yujin Tang ยท Sakana AI ยท arXiv
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LL...
Yejin Lee, Seungbeom Kim, Yongha Lee ... ยท Sungkyunkwan University ยท arXiv
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time sla...
Thursday, September 24, 2026
Rayhan Rashed, Senja Filipi, Ross Cutler ยท Microsoft +1 ยท arXiv
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested....
Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos ... ยท Purdue University +2 ยท arXiv
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-...
Alessandro Bondielli, Lucia Passaro, Serena Auriemma ... ยท University of Pisa ยท EMNLP 2026
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic informati...
Hongyao Deng, Wenhao Guan, Xuetao Lin ... ยท Xiamen University ยท arXiv
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Ed...