Audio ML Papers

Last 7 Days (October 02 - October 09, 2026)

Subcategories: All (20) | Speech Synthesis (0) | Music Synthesis (1) | Ambient Synthesis (3) | Quality Evaluation (0) | Enhancement (4) | Asr (0) | Llm Audio (5) | Midi Generation (0) | Generative Conditioning (1) | Other (6)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 81)
Yingda Shen, Yao Qian, Yuxuan Hu ... · University of Science and Technology of China +1 · arXiv
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framewo...
#2 TOP PAPER (Score: 81)
Sonal Kumar, Sinan Hersek, Artem Dementyev ... · Google DeepMind +4 · arXiv
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound locali...
#3 TOP PAPER (Score: 81)
William Chen, ★ Prem Seetharaman, Ke Chen, ★ Shinji Watanabe, ★ Justin Salamon ... · Adobe Research +1 · ICLR 2027 (inferred from footer "iclr2027_conference")
Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within o...
Thursday, October 08, 2026
Jing Xu, Cunjian Chen, Qiuhong Ke · Monash University · NeurIPS 2026
Group dance generation aims to synthesize coordinated multi-dancer choreography from music, with broad applications in animation and interactive content creation. This task requires modeling dense inter-person dependencies to ensure spatial coordination, while naturally preservin...
Jinbo Hu, Hang Su, Lichun Fan ... · Xiaomi Inc. · arXiv
Large audio-language models (LALMs) have achieved strong performance in general audio understanding, yet most are designed for monaural input and discard the inter-channel cues essential for spatial perception. In contrast, existing spatial audio-language models are purpose-built...
Aviad Dahan, Rajaei Khatib, Yonatan Bitton ... · Tel Aviv University +1 · arXiv
A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory. Rendering such a scene from a novel viewpoint requires that every source be available as an individual waveform, so that ...
Wednesday, October 07, 2026
William Chen, ★ Prem Seetharaman, Ke Chen, ★ Shinji Watanabe, ★ Justin Salamon ... · Adobe Research +1 · ICLR 2027 (inferred from footer "iclr2027_conference")
Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within o...
Pooneh Mousavi, ★ Mirco Ravanelli, Cem Subakan · Concordia University +2 · arXiv
Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is...
Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin ... · University of Virginia +2 · arXiv
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. W...
Tuesday, October 06, 2026
Pengjun Fang, Jingyi Fa, Kam Man Wu ... · The Hong Kong University of Science and Technology +3 · arXiv
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interac...
Kyudan Jung, Hyunsin Park, Yoonhyung Lee ... · Qualcomm AI Research +1 · ICLR 2027 (inferred from "iclr2027_conference" in text)
As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in rea...
Sungnyun Kim, Sungwoo Cho, Jihwan Oh ... · Korea Advanced Institute of Science and Technology (KAIST) +1 · arXiv
Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot reac...
Lee Seung-woo, Bowen Qi, Kim Min-jun ... · Pusan National University +2 · arXiv
Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose ...
Monday, October 05, 2026
Yingda Shen, Yao Qian, Yuxuan Hu ... · University of Science and Technology of China +1 · arXiv
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framewo...
Sunday, October 04, 2026
Sonal Kumar, Sinan Hersek, Artem Dementyev ... · Google DeepMind +4 · arXiv
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound locali...
Junyan Jiang, Ruibin Yuan, Jiahao Pan, ★ Yann LeCun ... · Unknown (likely Meta AI based on "m-a-p" HuggingFace handle and MERT lineage, but not explicitly stated in provided text) +1 · arXiv
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present Sheet...
Shuyuan Tu, Qi Tian, Yinming Huang ... · Tencent Hunyuan Foundation Model Team +2 · arXiv
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, ...
Team Kandinsky, Julia Agafonova, Bulat Akhmatov ... · GreenKandinsky Lab · arXiv
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 4...
Saturday, October 03, 2026
Séverin Baroudi, ★ Hervé Bredin, Ricard Marxer · Univ Toulon, Aix Marseille Univ, CNRS, LIS +6 · arXiv
Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction wit...
Ron Aluf, Alon Canfi, Eliya Nachmani · Ben-Gurion University of the Negev · NeurIPS 2026
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook ...
Xuanjun Chen, Zixiong Su, Hao Shi, ★ Hung-yi Lee ... · National Taiwan University +3 · arXiv
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate ...
Friday, October 02, 2026
Ping Wang, Guang Yang, Shao-Rong Su, ★ Noah A. Smith ... · University of Washington +1 · arXiv (Preprint)
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both...
David Braun, Adam Finkelstein · University of Illinois Urbana-Champaign · ISMIR 2026
This paper introduces PEACE, the first joint embedding of audio effect code and output audio. Building on SLAP's multimodal objective, we pair an AFx-Rep audio encoder with two code encoders for Faust, a functional language for audio signal processing. First, we evaluate a fine-t...