Audio ML Papers

Last 7 Days (July 16 - July 23, 2026)

Subcategories: All (24) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (1) | Quality Evaluation (0) | Enhancement (1) | Asr (1) | Llm Audio (0) | Midi Generation (1) | Generative Conditioning (0) | Other (20)
← Previous Week | Current Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 80)
David Ayllon, Alice Baird, Jeffrey Brooks ... ยท arXiv
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual represen...
#2 TOP PAPER (Score: 79)
Qiaoyu Yang, Lixing He, Binyue Deng ... ยท Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
#3 TOP PAPER (Score: 78)
Shuai Wang, Zihan Qian, Ke Zhang ... ยท IEEE SLT 2026
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the...
Wednesday, July 22, 2026
Siqian Tong, Xuan Li, Chaozhuo Li ... ยท arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
Rongshen He, Xinyu Liang, Dekun Chen ... ยท arXiv
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enab...
Tieyao Zhang, Yuke Liu, Jiaxing Yu ... ยท arXiv (preprint)
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning...
Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak ... ยท Interspeech 2026
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average...
Kaicheng Luo, Xuefei Gong, Yutao Sun ... ยท ASRU 2025
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignmen...
Tuesday, July 21, 2026
Laurin Wagner, Mario Zusag, Bernhard Thallinger ยท arXiv
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreli...
Yushan Yashengjiang, Jie Zhang, Miao Sun ... ยท arXiv (preprint)
Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, de...
Noah Schaffer, Nikhil Singh ยท arXiv
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement work...
Zhenglong Liu, Wangyou Zhang, Chenda Li ... ยท Interspeech 2026
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnosti...
Monday, July 20, 2026
Chen Yang, Ganye Wen, Bin Huang ... ยท arXiv
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling stat...
Yuxiang Zhao, Yichi Zhang, Yanjie An ... ยท arXiv (preprint)
Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, while proprietary products and AP...
Shengfan Shen, Di Wu, Xingchen Song ... ยท arXiv
Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightweight control layer that wraps around a TTS engine to externalize and govern its expressive behavior....
Sunday, July 19, 2026
Xiaoyu Yang, Xuenan Xu, Wenyi Yu ... ยท IEEE Transactions on Audio, Speech, and Language Processing (Submitted)
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether ...
Saturday, July 18, 2026
Qiaoyu Yang, Lixing He, Binyue Deng ... ยท Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
Yishan Lv, Jing Luo, Xinyu Yang ... ยท arXiv
Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length ...
Ye Lu, Yihan Yan, Zhaoyang Zhang ... ยท arXiv
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker chara...
Yu-Wen Chen, Julia Hirschberg ยท arXiv
Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited...
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern ... ยท IWAENC 2026
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a ...
Friday, July 17, 2026
Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen ... ยท ISMIR 2026
Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing appr...
Xin Wei, Shi He, Yihe Yuan ... ยท arXiv
In X-lingual automatic speaker verification (ASV), fixed front-end scores vary in reliability with language match, duration, and score source. We propose AMECxSV, an adaptive metadata-driven embedding-fusion calibration backend for metadata-available settings. AMECxSV fuses trial...
Thursday, July 16, 2026
David Ayllon, Alice Baird, Jeffrey Brooks ... ยท arXiv
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual represen...
Shuai Wang, Zihan Qian, Ke Zhang ... ยท IEEE SLT 2026
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the...
Akฤฑn Oktav ยท arXiv
Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than b...
Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir ... ยท arXiv
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) fo...