Audio ML Papers

Last 7 Days (August 12 - August 19, 2026)

Subcategories: All (20) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (1) | Midi Generation (0) | Generative Conditioning (0) | Other (19)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 75)
Eloi Moliner, Christoph Hold, Juan Azcarreta Ortiz ... ยท Reality Labs Research, Meta +4 ยท arXiv (preprint)
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spati...
#2 TOP PAPER (Score: 74)
Feng Yin, Shuai Shi, Junjie Zheng ... ยท VUI Labs Research ยท arXiv
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Qu...
#3 TOP PAPER (Score: 73)
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor ... ยท NVIDIA Corporation ยท arXiv
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing mu...
Tuesday, August 18, 2026
Feiyu Shen, Kun Xie, Yichen Wu, โ˜… Lei Xie ... ยท Xiaohongshu ยท arXiv
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlle...
Monday, August 17, 2026
Tony Alex, Wish Suharitdamrong, Sara Atito ... ยท University of Surrey +1 ยท arXiv
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of se...
Hanlin Zhang, Daxin Tan, Dehua Tao ... ยท City University of Hong Kong +2 ยท arXiv
Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training proce...
Fengji Ma, Yan Rong, Xu Li ... ยท Kling Team ยท arXiv
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, ...
Chen-An Li, โ˜… Hung-yi Lee ยท National Taiwan University +1 ยท arXiv
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including...
Tao Feng, Xu Li, Xiangyang Luo ... ยท Kuaishou Technology ยท arXiv
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily...
Sunday, August 16, 2026
Yusheng Dai, Kangdi Wang, Baolong Gao ... ยท Monash University +2 ยท ACM Multimedia 2026
Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on unc...
Nicholas Sanders, Gustav Eje Henter, Simon King ... ยท University of Edinburgh +3 ยท arXiv
Expressive text-to-speech (TTS) systems that use explicit conditioning labels provide direct and interpretable control over expressive attributes, in contrast to reference-based or prompting-based approaches, but require labeled data. Obtaining these labels at scale is costly and...
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko ... ยท Kandinsky Lab ยท arXiv
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized line...
Jiaming He, Zhicong Huang, Tian Jin ... ยท arXiv
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, wh...
Saturday, August 15, 2026
Ashish Anand Shukla, Rini Smita Thakur, Aryan Das ... ยท Indian Institute of Science Education and Research Bhopal +1 ยท ACM CIKM 2026
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annota...
Friday, August 14, 2026
Eloi Moliner, Christoph Hold, Juan Azcarreta Ortiz ... ยท Reality Labs Research, Meta +4 ยท arXiv (preprint)
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spati...
Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jimรฉnez, โ˜… Xavier Serra ... ยท Universitat Pompeu Fabra (Music Technology Group) +2 ยท ISMIR 2026
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better acr...
Thursday, August 13, 2026
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor ... ยท NVIDIA Corporation ยท arXiv
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing mu...
Wenxiang Guo, Changhao Pan, Ziyue Jiang ... ยท Zhejiang University ยท IEEE Transactions on Multimedia
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing. Existing Text-to-Audio (T2A) systems either reduce quoted speech to unintellig...
Wednesday, August 12, 2026
Feng Yin, Shuai Shi, Junjie Zheng ... ยท VUI Labs Research ยท arXiv
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Qu...
Jiabao Zhuang, Changhao Jiang, Hanchen Wang ... ยท Fudan University ยท arXiv
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, an...
Yining Wang ยท arXiv
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence ...
Xingwei Sun, Heinrich Dinkel, Gang Li ... ยท Xiaomi Inc. +3 ยท arXiv
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal opti...
Rong Chao, Sung-Feng Huang, Moreno La Quatra, โ˜… Yu Tsao ... ยท National Taiwan University +4 ยท INTERSPEECH 2026
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidt...