Audio ML Papers

Last 7 Days (September 27 - October 04, 2026)

Subcategories: All (23) | Speech Synthesis (4) | Music Synthesis (0) | Ambient Synthesis (2) | Quality Evaluation (0) | Enhancement (2) | Asr (1) | Llm Audio (5) | Midi Generation (0) | Generative Conditioning (0) | Other (9)
← Previous Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 84)
Ruibin Yuan, Jiahao Pan, Junyan Jiang ... · HKGAI · arXiv
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through...
#2 TOP PAPER (Score: 83)
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann · MIT CSAIL +2 · NeurIPS 2026
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound alon...
#3 TOP PAPER (Score: 81)
Yayue Deng, Dingdong Wang, Yuxuan Hu ... · The Chinese University of Hong Kong +1 · EMNLP 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target s...
Saturday, October 03, 2026
Séverin Baroudi, ★ Hervé Bredin, Ricard Marxer · Univ Toulon, Aix Marseille Univ, CNRS, LIS +6 · arXiv
Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction wit...
Ron Aluf, Alon Canfi, Eliya Nachmani · Ben-Gurion University of the Negev · NeurIPS 2026
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook ...
Xuanjun Chen, Zixiong Su, Hao Shi, ★ Hung-yi Lee ... · National Taiwan University +3 · arXiv
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate ...
Friday, October 02, 2026
Ping Wang, Guang Yang, Shao-Rong Su, ★ Noah A. Smith ... · University of Washington +1 · arXiv (Preprint)
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both...
David Braun, Adam Finkelstein · University of Illinois Urbana-Champaign · ISMIR 2026
This paper introduces PEACE, the first joint embedding of audio effect code and output audio. Building on SLAP's multimodal objective, we pair an AFx-Rep audio encoder with two code encoders for Faust, a functional language for audio signal processing. First, we evaluate a fine-t...
Thursday, October 01, 2026
Haibo Wang, Jiteng Mu, Jialu Li ... · Adobe Research +2 · arXiv
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence ac...
Yuxiang Wang, Kunyu Feng, Yuancheng Wang ... · Tencent Hunyuan +6 · ICLR 2027 (Preprint)
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays res...
Wednesday, September 30, 2026
Kyoungjun Park, Yunzhe Li, Lili Qiu · The University of Texas at Austin · arXiv
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their rank...
Rong Wan, Wei Xie, Jiaxi Li, ★ Wenwu Wang ... · University of Surrey +3 · arXiv
Audio deepfake detection (ADD) must remain effective when new spoofing attacks emerge after deployment. Emerging audio language model (ALM)-based ADD methods are built on predefined supervision from ground-truth labels or verified forensic rationales. However, this paradigm overl...
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka ... · NTT, Inc. +2 · Interspeech 2026
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance...
Shengbo Cai, Zhisheng Zhang, Zichao Nie ... · Tsinghua University · arXiv
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Expe...
Tuesday, September 29, 2026
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann · MIT CSAIL +2 · NeurIPS 2026
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound alon...
Yunzhe Li, Kyoungjun Park, Hongzi Zhu ... · The University of Texas at Austin +1 · arXiv
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and veri...
Duowen Chen, Jinjin He, Gouthaman KV ... · Georgia Institute of Technology +1 · NeurIPS 2026
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representat...
Jisoo Park, Seonghak Lee, Hyojin Park ... · Chung-Ang University +1 · NeurIPS 2026
Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fideli...
Lei Ke, Jiahao Pan, Zeyue Tian ... · Unknown (Affiliations are numbered 1,2 but names are not provided in the text) +1 · arXiv
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present...
Monday, September 28, 2026
Wenyi Yu, Siyin Wang, Terumi Chiba ... · Tsinghua University +1 · arXiv (Preprint)
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirement...
Tianxin Xie, Pengfei Zhang, Kai Jiang ... · The Hong Kong University of Science and Technology (Guangzhou) · arXiv
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for trai...
Donghang Wu, Haoyang Zhang, Yizhou Peng ... · Nanyang Technological University +2 · arXiv
Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like halluc...
Sunday, September 27, 2026
Ruibin Yuan, Jiahao Pan, Junyan Jiang ... · HKGAI · arXiv
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through...
Yayue Deng, Dingdong Wang, Yuxuan Hu ... · The Chinese University of Hong Kong +1 · EMNLP 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target s...
Yifan Yang, Xiaoyu Yang, Zengrui Jin, ★ Yuxuan Wang ... · Alibaba Group +6 · arXiv
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohi...
Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu · University of Zurich and ETH Zurich +4 · arXiv
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation...