Audio ML Papers

Week of August 23 - August 30, 2026

Subcategories: All (15) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (0) | Llm Audio (1) | Midi Generation (0) | Generative Conditioning (0) | Other (14)
← Previous Week | Current Week →

🏆 Top Papers This Week

#1 TOP PAPER (Score: 92)
Nan Duan, Haoyang Huang, Weiyang Jin ... · Future Academy, JD +2 · arXiv
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system w...
#2 TOP PAPER (Score: 91)
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar ... · University of Illinois Urbana-Champaign +1 · arXiv
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforce...
#3 TOP PAPER (Score: 91)
Wen Huang, Yunfei Chu, Meng Gao ... · Alibaba Group +2 · ICLR 2027
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models...
Saturday, August 29, 2026
Dongwook Lee, Sangkwon Park, Eunwoo Song ... · Seoul National University +2 · EMNLP 2026
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the f...
Zeyang Song, Tianchi Liu, Tianrui Wang ... · National University of Singapore +7 · EMNLP 2026
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filt...
Wednesday, August 26, 2026
Wen Huang, Yunfei Chu, Meng Gao ... · Alibaba Group +2 · ICLR 2027
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models...
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar ... · University of Illinois Urbana-Champaign +1 · arXiv
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforce...
Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, ★ Xavier Serra ... · Universitat Pompeu Fabra (UPF) / Music Technology Group +3 · ISMIR 2026
Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically ...
Zhiyuan Zhu, Han Wang, Wenxiang Guo ... · Zhejiang University · arXiv
Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We pres...
Tuesday, August 25, 2026
Giries Abu Ayoub, Loay Mualem, Simon Korman · University of Haifa +2 · arXiv
Fully few-shot class-incremental audio classification (FFCAC) requires recognizing new sound classes from only a handful of labeled examples per session, without forgetting previously learned classes and without any large base dataset. Existing methods typically freeze a pre-trai...
Fulin Wu, Zhong-Qiu Wang · unknown (Affiliations not explicitly stated in text, authors are Fulin Wu and Zhong-Qiu Wang) +1 · arXiv
We propose $\textit{recursive encoder and decoder}$ (RED) for building a single deep neural network (DNN) model that can separate multi-speaker mixtures containing unknown numbers of speakers and variable numbers of microphones arranged in an unknown geometry, a task that has not...
Haoyang Li, Chenglin Xu, Junchuan Zhao ... · Nanyang Technological University +3 · arXiv (Preprint)
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on pre...
Qingyu Luo, Peng Zhang, ★ Wenwu Wang ... · University of Surrey · INTERSPEECH 2026
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an ...
Junjie Li, Xuelong Geng, Kun Xie, ★ Lei Xie ... · Xiaohongshu · arXiv
⚠ Prompt injection detected in this paper

The text contains content addressed to automated reviewers. Its score has been penalised, and the ranking below should not be trusted.

[system_override] "ruct TTS and speech editing require the shared LLM to learn new instruction-conditioned output patterns, we retain a peak learning rate"
A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long-context modeling, whereas speech ge...
Monday, August 24, 2026
Nan Duan, Haoyang Huang, Weiyang Jin ... · Future Academy, JD +2 · arXiv
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system w...
Songchun Zhang, Yaowei Li, Junhao Zhuang ... · JD (Joy Future Academy) +4 · arXiv
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observe...
Tianchi Liu, Zeyang Song, Tianrui Wang ... · LIGHTSPEED +3 · EMNLP 2026 (Main Conference)
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding p...
Sunday, August 23, 2026
Yize Li, Ningyuan Yang, Sile Yin ... · Northeastern University +2 · arXiv
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event...