Audio ML Papers

Last 7 Days (September 30 - October 07, 2026)

Subcategories: All (20) | Speech Synthesis (1) | Music Synthesis (1) | Ambient Synthesis (2) | Quality Evaluation (0) | Enhancement (3) | Asr (0) | Llm Audio (7) | Midi Generation (0) | Generative Conditioning (0) | Other (6)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 81)
Kyoungjun Park, Yunzhe Li, Lili Qiu ยท The University of Texas at Austin ยท arXiv
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their rank...
#2 TOP PAPER (Score: 81)
Yingda Shen, Yao Qian, Yuxuan Hu ... ยท University of Science and Technology of China +1 ยท arXiv
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framewo...
#3 TOP PAPER (Score: 81)
Sonal Kumar, Sinan Hersek, Artem Dementyev ... ยท Google DeepMind +4 ยท arXiv
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound locali...
Tuesday, October 06, 2026
Kyudan Jung, Hyunsin Park, Yoonhyung Lee ... ยท Qualcomm AI Research +1 ยท ICLR 2027 (inferred from "iclr2027_conference" in text)
As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in rea...
Sungnyun Kim, Sungwoo Cho, Jihwan Oh ... ยท Korea Advanced Institute of Science and Technology (KAIST) +1 ยท arXiv
Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot reac...
Lee Seung-woo, Bowen Qi, Kim Min-jun ... ยท Pusan National University +2 ยท arXiv
Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose ...
Monday, October 05, 2026
Yingda Shen, Yao Qian, Yuxuan Hu ... ยท University of Science and Technology of China +1 ยท arXiv
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framewo...
Sunday, October 04, 2026
Sonal Kumar, Sinan Hersek, Artem Dementyev ... ยท Google DeepMind +4 ยท arXiv
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound locali...
Junyan Jiang, Ruibin Yuan, Jiahao Pan, โ˜… Yann LeCun ... ยท Unknown (likely Meta AI based on "m-a-p" HuggingFace handle and MERT lineage, but not explicitly stated in provided text) +1 ยท arXiv
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present Sheet...
Shuyuan Tu, Qi Tian, Yinming Huang ... ยท Tencent Hunyuan Foundation Model Team +2 ยท arXiv
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, ...
Team Kandinsky, Julia Agafonova, Bulat Akhmatov ... ยท GreenKandinsky Lab ยท arXiv
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 4...
Saturday, October 03, 2026
Sรฉverin Baroudi, โ˜… Hervรฉ Bredin, Ricard Marxer ยท Univ Toulon, Aix Marseille Univ, CNRS, LIS +6 ยท arXiv
Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction wit...
Ron Aluf, Alon Canfi, Eliya Nachmani ยท Ben-Gurion University of the Negev ยท NeurIPS 2026
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook ...
Xuanjun Chen, Zixiong Su, Hao Shi, โ˜… Hung-yi Lee ... ยท National Taiwan University +3 ยท arXiv
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate ...
Friday, October 02, 2026
Ping Wang, Guang Yang, Shao-Rong Su, โ˜… Noah A. Smith ... ยท University of Washington +1 ยท arXiv (Preprint)
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both...
David Braun, Adam Finkelstein ยท University of Illinois Urbana-Champaign ยท ISMIR 2026
This paper introduces PEACE, the first joint embedding of audio effect code and output audio. Building on SLAP's multimodal objective, we pair an AFx-Rep audio encoder with two code encoders for Faust, a functional language for audio signal processing. First, we evaluate a fine-t...
Thursday, October 01, 2026
Haibo Wang, Jiteng Mu, Jialu Li ... ยท Adobe Research +2 ยท arXiv
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence ac...
Yuxiang Wang, Kunyu Feng, Yuancheng Wang ... ยท Tencent Hunyuan +6 ยท ICLR 2027 (Preprint)
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays res...
Wednesday, September 30, 2026
Kyoungjun Park, Yunzhe Li, Lili Qiu ยท The University of Texas at Austin ยท arXiv
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their rank...
Rong Wan, Wei Xie, Jiaxi Li, โ˜… Wenwu Wang ... ยท University of Surrey +3 ยท arXiv
Audio deepfake detection (ADD) must remain effective when new spoofing attacks emerge after deployment. Emerging audio language model (ALM)-based ADD methods are built on predefined supervision from ground-truth labels or verified forensic rationales. However, this paradigm overl...
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka ... ยท NTT, Inc. +2 ยท Interspeech 2026
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance...
Shengbo Cai, Zhisheng Zhang, Zichao Nie ... ยท Tsinghua University ยท arXiv
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Expe...
Junming Lin, โ˜… Yuxuan Wang, Zhenxin Lei ... ยท arXiv
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluat...