Audio ML Papers

Last 7 Days (July 17 - July 24, 2026)

Subcategories: All (23) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (1) | Quality Evaluation (0) | Enhancement (1) | Asr (1) | Llm Audio (0) | Midi Generation (1) | Generative Conditioning (0) | Other (19)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 79)
Qiaoyu Yang, Lixing He, Binyue Deng ... · Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
#2 TOP PAPER (Score: 74)
Siqian Tong, Xuan Li, Chaozhuo Li ... · arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
#3 TOP PAPER (Score: 73)
Xinjie Zhang, Peng Zhang, Shicheng Zheng ... · arXiv
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed compo...
Thursday, July 23, 2026
Songchen Xu, Ting Song, Shaohan Huang ... · arXiv
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization ...
Shengkui Zhao, Zexu Pan, Haoxu Wang ... · arXiv (preprint)
Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range ...
Viola Negroni, Xin Wang, Wanying Ge ... · arXiv
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is ...
Wednesday, July 22, 2026
Siqian Tong, Xuan Li, Chaozhuo Li ... · arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
Rongshen He, Xinyu Liang, Dekun Chen ... · arXiv
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enab...
Tieyao Zhang, Yuke Liu, Jiaxing Yu ... · arXiv (preprint)
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning...
Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak ... · Interspeech 2026
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average...
Tuesday, July 21, 2026
Xinjie Zhang, Peng Zhang, Shicheng Zheng ... · arXiv
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed compo...
Laurin Wagner, Mario Zusag, Bernhard Thallinger · arXiv
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreli...
Yushan Yashengjiang, Jie Zhang, Miao Sun ... · arXiv (preprint)
Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, de...
Noah Schaffer, Nikhil Singh · arXiv
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement work...
Haolin He, Renhe Sun, Zheqi Dai ... · DCASE 2026 (IEEE Workshop on Detection and Classification of Acoustic Scenes and Events)
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a larg...
Zhenglong Liu, Wangyou Zhang, Chenda Li ... · Interspeech 2026
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnosti...
Monday, July 20, 2026
Chen Yang, Ganye Wen, Bin Huang ... · arXiv
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling stat...
Yuxiang Zhao, Yichi Zhang, Yanjie An ... · arXiv (preprint)
Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, while proprietary products and AP...
Sunday, July 19, 2026
Xiaoyu Yang, Xuenan Xu, Wenyi Yu ... · IEEE Transactions on Audio, Speech, and Language Processing (Submitted)
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether ...
Saturday, July 18, 2026
Qiaoyu Yang, Lixing He, Binyue Deng ... · Interspeech 2026
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio ...
Yishan Lv, Jing Luo, Xinyu Yang ... · arXiv
Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length ...
Ye Lu, Yihan Yan, Zhaoyang Zhang ... · arXiv
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker chara...
Yu-Wen Chen, Julia Hirschberg · arXiv
Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited...
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern ... · IWAENC 2026
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a ...
Friday, July 17, 2026
Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen ... · ISMIR 2026
Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing appr...
Xin Wei, Shi He, Yihe Yuan ... · arXiv
In X-lingual automatic speaker verification (ASV), fixed front-end scores vary in reliability with language match, duration, and score source. We propose AMECxSV, an adaptive metadata-driven embedding-fusion calibration backend for metadata-available settings. AMECxSV fuses trial...