Audio ML Papers

Last 7 Days (July 19 - July 26, 2026)

Subcategories: All (23) | Speech Synthesis (0) | Music Synthesis (0) | Ambient Synthesis (1) | Quality Evaluation (0) | Enhancement (1) | Asr (1) | Llm Audio (0) | Midi Generation (1) | Generative Conditioning (0) | Other (19)
← Previous Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 75)
Yunjin Gu, Qianrui Zhou, Hua Xu ยท ACM Multimedia 2026 (MM '26)
Unsupervised multimodal intent discovery aims to uncover latent intents from unlabeled multimodal dialogues, but remains challenging due to the lack of explicit semantic supervision. Existing methods often provide limited interpretability, as their refinement mainly relies on geo...
#2 TOP PAPER (Score: 74)
Siqian Tong, Xuan Li, Chaozhuo Li ... ยท arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
#3 TOP PAPER (Score: 73)
Xinjie Zhang, Peng Zhang, Shicheng Zheng ... ยท arXiv
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed compo...
Saturday, July 25, 2026
Shreshth Saini, Neil Birkbeck, Yilin Wang ... ยท arXiv (Preprint)
Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching makes each rollout 2-3x faster at near-lossless quality. Composition is safe only if lossy caching pre...
Friday, July 24, 2026
Yunjin Gu, Qianrui Zhou, Hua Xu ยท ACM Multimedia 2026 (MM '26)
Unsupervised multimodal intent discovery aims to uncover latent intents from unlabeled multimodal dialogues, but remains challenging due to the lack of explicit semantic supervision. Existing methods often provide limited interpretability, as their refinement mainly relies on geo...
Ziyu Wang, Kun Fang, Yann LeCun ยท arXiv (Submitted to ISMIR 2025 based on template reference)
Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remain...
Yining Yang, Ruogu Chen, Jie Han ยท International Society for Music Information Retrieval (ISMIR)
Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction ...
Pengfei Zhang, Biao Tian, Tianxin Xie ... ยท arXiv
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text le...
Thursday, July 23, 2026
Daniyal Kabir Dar, Arun Ross ยท IEEE/IAPR International Joint Conference on Biometrics (IJCB) 2026
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different d...
Songchen Xu, Ting Song, Shaohan Huang ... ยท arXiv
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization ...
Shengkui Zhao, Zexu Pan, Haoxu Wang ... ยท arXiv (preprint)
Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range ...
Hao Zhang, Yiwen Zhao, Yixuan Zhang ... ยท arXiv (preprint)
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based ag...
Viola Negroni, Xin Wang, Wanying Ge ... ยท arXiv
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is ...
Wednesday, July 22, 2026
Siqian Tong, Xuan Li, Chaozhuo Li ... ยท arXiv (Preprint)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or...
Rongshen He, Xinyu Liang, Dekun Chen ... ยท arXiv
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enab...
Tieyao Zhang, Yuke Liu, Jiaxing Yu ... ยท arXiv (preprint)
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning...
Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak ... ยท Interspeech 2026
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average...
Tuesday, July 21, 2026
Xinjie Zhang, Peng Zhang, Shicheng Zheng ... ยท arXiv
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed compo...
Laurin Wagner, Mario Zusag, Bernhard Thallinger ยท arXiv
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreli...
Yushan Yashengjiang, Jie Zhang, Miao Sun ... ยท arXiv (preprint)
Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, de...
Noah Schaffer, Nikhil Singh ยท arXiv
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement work...
Haolin He, Renhe Sun, Zheqi Dai ... ยท DCASE 2026 (IEEE Workshop on Detection and Classification of Acoustic Scenes and Events)
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a larg...
Zhenglong Liu, Wangyou Zhang, Chenda Li ... ยท Interspeech 2026
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnosti...
Monday, July 20, 2026
Chen Yang, Ganye Wen, Bin Huang ... ยท arXiv
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling stat...
Yuxiang Zhao, Yichi Zhang, Yanjie An ... ยท arXiv (preprint)
Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, while proprietary products and AP...
Sunday, July 19, 2026
Xiaoyu Yang, Xuenan Xu, Wenyi Yu ... ยท IEEE Transactions on Audio, Speech, and Language Processing (Submitted)
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether ...