Audio ML Papers

Week of July 12 - July 19, 2026

Subcategories: All (23) | Speech Synthesis (0) | Music Synthesis (1) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (0) | Asr (1) | Llm Audio (0) | Midi Generation (0) | Generative Conditioning (0) | Other (21)
← Previous Week | Current Week →

🏆 Top Papers This Week

#1 TOP PAPER (Score: 81)
Jin Xu, Kangdi Wang, Ruibin Yuan ... · arXiv
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions...
#2 TOP PAPER (Score: 80)
David Ayllon, Alice Baird, Jeffrey Brooks ... · arXiv
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual represen...
#3 TOP PAPER (Score: 79)
Qiaoyu Yang, Lixing He, Binyue Deng ... · Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
Saturday, July 18, 2026
Qiaoyu Yang, Lixing He, Binyue Deng ... · Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
Yishan Lv, Jing Luo, Xinyu Yang ... · arXiv
Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length ...
Ye Lu, Yihan Yan, Zhaoyang Zhang ... · arXiv
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker chara...
Yu-Wen Chen, Julia Hirschberg · arXiv
Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited...
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern ... · IWAENC 2026
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a ...
Friday, July 17, 2026
Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen ... · ISMIR 2026
Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing appr...
Xin Wei, Shi He, Yihe Yuan ... · arXiv
In X-lingual automatic speaker verification (ASV), fixed front-end scores vary in reliability with language match, duration, and score source. We propose AMECxSV, an adaptive metadata-driven embedding-fusion calibration backend for metadata-available settings. AMECxSV fuses trial...
Thursday, July 16, 2026
David Ayllon, Alice Baird, Jeffrey Brooks ... · arXiv
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual represen...
Shuai Wang, Zihan Qian, Ke Zhang ... · IEEE SLT 2026
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the...
Akın Oktav · arXiv
Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than b...
Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir ... · arXiv
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) fo...
Wednesday, July 15, 2026
Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth ... · ISMIR 2026
Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such as full-score prediction, and therefore do not match the broader range of operations that arise in analysis workflows, including partial com...
Wangjin Zhou, Yizhou Zhang, Yichi Wang ... · arXiv
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceil...
Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim ... · Interspeech 2026
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limit...
Tuesday, July 14, 2026
Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani ... · arXiv
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-...
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes ... · International Conference on Digital Audio Effects (DAFx) 2026
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio typ...
Monday, July 13, 2026
Jin Xu, Kangdi Wang, Ruibin Yuan ... · arXiv
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions...
Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu ... · arXiv
Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective inference-time intervention, ...
Zhen-Lin Chen, Maosen Sheng, Peng Lin ... · SIGIR 2026
Multimodal information is pivotal for e-commerce search ranking. Existing works leverage multimodal data typically by fine-tuning general Multimodal Large Language Models (MLLMs) via collaborative signals, subsequently integrating the derived representations into ranking models a...
Mingyue Huo, Yuheng Zhang, Hao Zhang · arXiv
Speech enhancement (SE) can substantially improve perceptual quality, yet enhanced speech does not necessarily improve automatic speech recognition (ASR). Existing remedies, such as retraining the enhancer jointly with recognizer or interpolating enhanced speech with the noisy in...
Chong Jing, Junan Zhang, Jing Yang ... · arXiv
Zero-shot instrument cloning aims to render an arbitrary [Target MIDI] sequence with the acoustic identity of an unseen instrument given only a short [Reference Audio, Reference MIDI] pair. Existing methods rely on pre-trained embeddings (e.g., CLAP) that compress the reference a...
Sunday, July 12, 2026
Stefano Bannò, Penny Karanasou, Mengjie Qian ... · arXiv
Automated assessment of second language (L2) speaking proficiency relies on large-scale annotated speech data, which remains scarce compared to widely available written learner corpora. A promising direction for addressing this imbalance is to use text-to-speech (TTS) and voice c...
Ryota Kimura, Sangheon Park, Natalia Polouliakh ... · arXiv
Dance-to-music generation is a promising task for applications such as choreography support and automatic accompaniment, where temporal coordination between body movement and sound is essential. In particular, using human joint positions as the motion representation is attractive...