We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spatial capture capabilities, such as irregular or sparse arrays, as classical linear methods fail to reconstruct high-order spatial detail. We introduce a diffusion-based generative framework that models the statistical properties of HOA RIRs. This enables device-agnostic encoding from arbitrary microphone arrays, potentially unseen during data measurement. Our approach incorporates a posterior sampling procedure that enforces consistency between the estimated signals and the measurements while plausibly reconstructing spatial information that is unobservable from the limited measurements alone. Experiments on simulated data demonstrate that our method outperforms linear and neural baselines, achieving accurate HOA RIR estimation up to 12th order. A listening test with binaural renderings, including both simulated and measured RIRs, further confirms that the proposed method yields higher perceptual similarity to reference Ambisonics RIRs than all baselines. The flexibility and accuracy of the proposed framework opens new possibilities for scalable acoustics simulations.
Primary: Reality Labs Research, Meta
All Institutions: Reality Labs Research, Meta, Aalto University, Acoustics Lab
This paper presents a significant advancement in spatial audio processing by effectively leveraging diffusion models to solve the ill-posed problem of high-order Ambisonics encoding from sparse measurements, achieving state-of-the-art perceptual and objective performance.
The paper proposes a novel application of diffusion models to the ill-posed inverse problem of encoding Room Impulse Responses (RIRs) into High-Order Ambisonics (HOA) from sparse or irregular microphone arrays. The core methodological contribution is the integration of a device-agnostic diffusion prior with a posterior sampling procedure that enforces data consistency via a range-projected likelihood term. The use of a hybrid time-frequency/time-domain architecture (NCSN++ backbone with a U-Net refinement stage for early reflections) is a well-reasoned design choice given the distinct temporal characteristics of RIRs. The introduction of a compressed spectrogram distance metric for the likelihood guidance is a technically sound innovation to emphasize weak high-order components. The formulation separates the device-agnostic prior from the device-specific likelihood, which is a strong theoretical foundation for generalization.
The experimental evaluation is rigorous and comprehensive. The authors utilize two datasets: a large-scale internal FDTD simulation dataset and the public Treble-10 dataset. They compare against three strong baselines: Linear Least-Squares, a Time-Dependent Neural Encoder, and a Conditional Diffusion model specific to the device. The inclusion of both objective metrics (EDC, NPM) and a subjective listening test (MUSHRA-style with binaural rendering) provides robust validation. The results demonstrate clear superiority in both perceptual similarity and objective metrics, particularly in preserving early reflections and late reverberation tails. The ablation studies effectively isolate the contribution of the range-projection constraint and the hybrid architecture.
The paper provides sufficient detail for reproduction, including dataset descriptions, model architectures (NCSN++ based), training hyperparameters (AdamW, learning rate, batch size), and specific implementation details like the compressed spectrogram definition. The use of standard libraries (PyTorch) and well-known architectures aids reproducibility. However, the reliance on an internal FDTD dataset for the primary training results limits independent verification of the scale of the results, though the public Treble-10 results offer some ground truth.
The primary limitation is the distribution gap between simulated training data and real-world measured RIRs, which the authors acknowledge and demonstrate leads to performance drops on measured data (Eigenmike-64). The method assumes a highly accurate Array Transfer Function (ATF), which may not hold in real-world deployments with calibration errors. Additionally, the iterative sampling process is computationally intensive, making it less suitable for real-time applications without significant acceleration techniques. The evaluation is currently limited to a single device configuration (Aria Glasses) for the device-agnostic claim, though the method is theoretically general.
This work has significant implications for spatial audio processing, particularly for Virtual Reality (VR) and Augmented Reality (AR) applications where scalable and accurate acoustic simulation is crucial. By enabling high-quality HOA encoding from sparse arrays, it lowers the barrier for capturing spatial audio with wearable devices. It also facilitates the generation of large-scale training data for spatial audio models. The device-agnostic nature of the approach promotes interoperability across different hardware platforms. This paper presents a significant advancement in spatial audio processing by effectively leveraging diffusion models to solve the ill-posed problem of high-order Ambisonics encoding from sparse measurements, achieving state-of-the-art perceptual and objective performance.
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.
Primary: VUI Labs Research
All Institutions: VUI Labs Research
Luna-TTS introduces a scalable, block-diffusion-based TTS framework that achieves state-of-the-art quality and latency by adapting pretrained LLMs through progressive architectural changes and RL post-training, effectively resolving the latency-quality trade-off inherent in previous diffusion TTS systems.
The paper proposes a novel architecture for Text-to-Speech (TTS) by adapting Large Language Models (LLMs) into Diffusion Language Models (dLLMs). Specifically, it introduces a "block-causal" attention mechanism that allows for streaming generation while retaining the parallel denoising benefits of diffusion models. The approach involves a progressive adaptation from causal to bidirectional and finally to block-causal attention on a pretrained AR text LLM. A key technical contribution is the application of Group Relative Policy Optimization (GRPO) directly over the realized denoising trajectory, optimizing for content correctness and speaker similarity. The use of a semantic-distilled RVQ tokenizer where the first codebook is anchored to linguistic content is a significant methodological detail that facilitates the diffusion process. The transition from fully parallel masked diffusion to block-autoregressive diffusion for latency reduction is a well-reasoned engineering solution to a known bottleneck in diffusion-based generation.
The evaluation is comprehensive, covering both objective metrics (CER, WER, SIM) and subjective/human-rated metrics for emotion and non-verbal vocalization (NVV) control. The paper claims state-of-the-art performance on Seed-TTS-Eval and CV3-Eval, outperforming both open-source and leading commercial systems. The inclusion of a "warmed serving protocol" to measure real-time factor (RTF) and first-block latency adds practical relevance to the experimental setup. The comparison against commercial baselines strengthens the claim of production-readiness. However, the specific numerical results are redacted in the provided text, preventing a precise verification of the magnitude of improvement, though the qualitative claims are strong.
The paper provides detailed descriptions of the architecture, including the tokenizer design, the masked diffusion formulation, and the training schedule. The mention of "0.6B backbone lineage" and "1 million hours of speech" provides scale context. However, the specific hyperparameters for the GRPO stage and the exact implementation details of the block-causal attention masking are likely critical for reproduction and may be subject to interpretation. The lack of publicly released code or weights (implied by "none" for URLs) significantly hinders immediate reproducibility for the broader community, although the technical report format suggests a high level of detail.
The primary limitation is the reliance on a specific RVQ tokenizer and the assumption that the semantic anchoring of the first codebook generalizes well across all languages and domains. The block-diffusion approach, while reducing latency, still requires multiple denoising steps per block, which may not be as fast as single-step AR decoding in extreme low-latency scenarios. Furthermore, the complexity of training a diffusion LLM with RL post-training is significantly higher than standard AR TTS training, potentially limiting accessibility. The paper does not extensively discuss the failure modes of the NVV control or the robustness of the emotion conditioning under adversarial text inputs.
This work represents a significant step towards making diffusion-based generative models viable for real-time, production-grade audio applications. By bridging the gap between the quality of diffusion models and the latency requirements of streaming TTS, it expands the toolkit available for developers building voice assistants, gaming NPCs, and accessibility tools. The emphasis on expressive control (emotion, NVVs) aligns with the growing demand for more natural and engaging human-computer interaction. However, the potential for misuse in deepfake generation remains a concern, necessitating robust safety measures in deployment. Luna-TTS introduces a scalable, block-diffusion-based TTS framework that achieves state-of-the-art quality and latency by adapting pretrained LLMs through progressive architectural changes and RL post-training, effectively resolving the latency-quality trade-off inherent in previous diffusion TTS systems.
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.
Primary: The Chinese University of Hong Kong, Shenzhen
All Institutions: The Chinese University of Hong Kong, Shenzhen, LIGHTSPEED, Independent Researcher
[One sentence main contribution]. The paper introduces Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and video responses by decoupling visual intent planning (VTP) and acoustic timing (speech units) to enable efficient training from heterogeneous data sources. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The work is a significant engineering achievement in integrating large language models with speech and video generation for interactive dialogue. Its primary value lies in the system-level design: the VTP interface and the shared acoustic-temporal interface allow for modular training and inference. While not introducing a new foundational model, it effectively demonstrates how to overcome data scarcity in omni-modal dialogue by leveraging structured intermediate representations. The "Prefix Streaming" technique is a notable contribution to the specific problem of autoregressive video generation stability. The paper is well-written, thoroughly evaluated, and addresses a timely and important problem in multimodal AI.
The paper proposes Ex-Omni-2D, a framework for omni-modal dialogue that generates text, speech, and video. The core methodological contribution is the "Visual Thought Plan" (VTP), a structured textual representation of visual intent (scene, emotion, motion) generated by the LLM, and a "native multi-codebook speech unit" interface. The architecture factors the problem: the LLM generates VTP and text; a Speech Generator produces acoustic units; and a Video Generator (based on Wan2.1) synthesizes video conditioned on the reference image, VTP, and frame-aligned speech units. A key technical detail is the "Prefix Streaming" mechanism for the video generator, which attempts to mitigate cumulative degradation in autoregressive video generation by carrying a clean latent from the previous chunk. The approach is a sophisticated integration of existing components (LLM, TTS, Video Diffusion) rather than a fundamental architectural breakthrough in any single modality. The factorization to avoid paired dialogue-video data is a pragmatic engineering solution to a data scarcity problem.
The evaluation covers audio quality (PQ, CU, SIM), video quality (SC, IQ, DD), and synchronization (Sync-C). The paper demonstrates that the Teacher model achieves high quality but is slow, while the Streaming Student offers a trade-off. The ablation studies on VTP and speech conditioning are valuable, showing that the VTP improves subject consistency and synchronization compared to a neutral plan. However, the video quality metrics (SC ~94%, IQ ~67%) are competitive but not state-of-the-art for dedicated video generators, which is expected given the constraints of dialogue-native generation. The audio quality is strong, leveraging recent TTS advances. The evaluation protocol is rigorous, using established benchmarks like VoiceBench and OmniCharacter.
The paper provides detailed descriptions of the architecture, training stages, and hyperparameters. The use of open-source backbones (Qwen3, Wan2.1) facilitates reproduction. The specific "Prefix Streaming" mechanism and data construction pipelines are described in sufficient detail for replication. The project page link suggests code or at least detailed visualizations may be available, though the code repository itself is not explicitly linked in the text provided.
The authors acknowledge several limitations: the speaker similarity is not perfect, the VTP is not independently sufficient for video control (requiring acoustic conditioning), and there is a measurable trade-off between language capability and visual planning in the shared LLM channel. The Streaming Student still has significant latency (first video chunk >3s) and is not real-time. The quality-efficiency trade-off means the fast model is significantly lower quality than the teacher.
This work contributes to the field of embodied AI and virtual humans by providing a framework for generating coherent, multi-modal responses. It addresses the challenge of generating visually present dialogue agents, which has applications in customer service, education, and entertainment. The factorization strategy offers a blueprint for training complex multi-modal systems where paired data is scarce. [One sentence main contribution]. The paper introduces Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and video responses by decoupling visual intent planning (VTP) and acoustic timing (speech units) to enable efficient training from heterogeneous data sources. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The work is a significant engineering achievement in integrating large language models with speech and video generation for interactive dialogue. Its primary value lies in the system-level design: the VTP interface and the shared acoustic-temporal interface allow for modular training and inference. While not introducing a new foundational model, it effectively demonstrates how to overcome data scarcity in omni-modal dialogue by leveraging structured intermediate representations. The "Prefix Streaming" technique is a notable contribution to the specific problem of autoregressive video generation stability. The paper is well-written, thoroughly evaluated, and addresses a timely and important problem in multimodal AI.
Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal alignment and speaker-utterance ambiguity in multi-speaker settings. To overcome these limitations, we propose CineDub, a unified diffusion-based model that achieves precise multi-speaker dialogue dubbing directly from uncropped videos, without face cropping or speaker diarization. Central to our approach is the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm, where holistic visual representations and a semantic-bundled transcription format are encoded independently, yet implicitly coupled through cross-modal training to resolve speaker ambiguity and enable precise multi-speaker multi-turn dialogue dubbing. Building on the unified temporal cues captured by holistic visual features, we further extend CineDub to joint speech and audio generation. We introduce an Ambient-to-Linguistic Curriculum Learning (ALC) to mitigate sub-task degradation, and a decoupled textual branch control mechanism to resolve cross-prompt interference during simultaneous generation. We also release two in-the-wild benchmarks, CineDub-Multi for multi-speaker dialogue dubbing and CineDub-SA for video-to-speech-and-audio (V2SA) generation, to enable evaluation under realistic conditions. Experiments show that CineDub achieves state-of-the-art results on established single-speaker dubbing and video-to-audio benchmarks while excelling in multi-speaker dialogue dubbing and acoustically coherent joint generation.
Primary: Monash University
All Institutions: Monash University, University of Chinese Academy of Sciences, Tsinghua University
CineDub presents a robust and innovative approach to multi-speaker video dubbing by effectively decoupling holistic visual conditioning from semantic transcription, achieving state-of-the-art results and providing valuable new benchmarks for the community.
The paper proposes CineDub, a unified diffusion-based framework for end-to-end video dubbing that operates on uncropped videos, addressing the limitations of hierarchical methods (which rely on brittle preprocessing like face cropping and diarization) and holistic methods (which suffer from speaker-utterance ambiguity). The core technical contribution is the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm. This involves using SynchFormer features as a holistic visual condition, which the authors argue captures both event-level audio-visual correspondence and fine-grained lip-sync cues through emergent attention mechanisms. To resolve speaker ambiguity, they introduce a "semantic-bundled transcription" format, where speaker descriptions are coupled with transcript segments, encoded by a pre-trained LLM (Gemma-T5). The framework extends to joint speech and audio generation (V2SA), introducing two key mechanisms: Ambient-to-Linguistic Curriculum Learning (ALC) to mitigate sub-task degradation by training on audio first, then speech, and a decoupled textual branch control mechanism to prevent cross-prompt interference between speech and audio conditions in the diffusion transformer. The approach is technically sound, leveraging recent advances in diffusion transformers and multimodal alignment, but the novelty lies primarily in the specific conditioning strategies and curriculum learning design rather than fundamental architectural changes.
The authors evaluate CineDub on single-speaker dubbing (GRID, CHEM), multi-speaker dubbing (new CineDub-Multi benchmark), video-to-audio (VGGSound), and joint V2SA (new CineDub-SA benchmark). Results show CineDub outperforms hierarchical baselines (like HPMDubbing, Speak2Dub) and holistic baselines (DeepDubber, DeepAudio) on most metrics, particularly in zero-shot voice cloning and multi-speaker scenarios. The introduction of two new benchmarks, CineDub-Multi and CineDub-SA, is a significant contribution, addressing the lack of realistic, in-the-wild evaluation data for multi-speaker and joint generation tasks. The ablation studies effectively demonstrate the necessity of the semantic-bundled transcription, the ALC curriculum, and the decoupled branches. The use of standard metrics (WER, LSE-D, UTMOS, FDVGG) is appropriate, though the reliance on embedding-based metrics for audio quality (FDVGG, KL) has known limitations regarding perceptual fidelity, which the authors partially address with UTMOS.
The paper provides detailed descriptions of the model architecture, training stages, and data processing pipelines. It mentions the use of pre-trained models (SynchFormer, Gemma-T5, CLIP) and specific datasets (VGGSound, AudioSet, SpeakerVid-5M). The release of the CineDub-Multi and CineDub-SA benchmarks enhances reproducibility for future work. However, the code is not explicitly linked in the provided text (only a demo page is mentioned), and the specific hyperparameters for the curriculum learning and meta-token initialization are not fully detailed in the excerpt, which might hinder exact replication.
The paper acknowledges that SynchFormer's attention-switching is not perfectly reliable, occasionally drifting or oscillating, which the semantic-bundled transcription aims to mitigate but may not fully resolve in all edge cases (e.g., heavy occlusion or off-screen speech). The reliance on MLLMs (Gemini 2.5 Pro) for generating the semantic-bundled transcriptions introduces a dependency on external models and potential errors in annotation, although manual verification is claimed. The joint generation model, while competitive, may still suffer from subtle acoustic incoherence compared to specialized single-task models, as indicated by some metric gaps. The benchmarks, while novel, are limited in size (139 and 562 samples respectively), which may not fully capture the diversity of in-the-wild scenarios.
CineDub has significant potential for multimedia production, enabling scalable and realistic video dubbing for movies, TV shows, and online content, particularly in multi-speaker settings. The release of new benchmarks will facilitate further research in this area. However, the technology also raises concerns about deepfake creation and misuse in generating deceptive audio-visual content. The authors' emphasis on realistic evaluation and the complexity of the pipeline may act as a slight barrier to malicious use, but the dual-use nature of such generative models remains a concern. CineDub presents a robust and innovative approach to multi-speaker video dubbing by effectively decoupling holistic visual conditioning from semantic transcription, achieving state-of-the-art results and providing valuable new benchmarks for the community.
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time. The reference is injected through two complementary signals: its diffusion latents are prepended to the audio stream, and a global speaker embedding modulates token of the target audio. On a benchmark of 674 speaker-text pairs spanning 30 speakers we compare against five strong voice-cloning text-to-speech baselines: our enhanced 5B model attains the highest speaker-encoder cosine similarity (SECS) across three independent verification networks (ECAPA-TDNN, WavLM-SV, Resemblyzer), statistically significantly outperforming every baseline. A side product of the architecture is that the audio path can be evaluated without the video path at inference time, yielding a ~30x speed-up over the full audio-video diffusion loop while preserving the voice-cloning behaviour.
Primary: Kandinsky Lab
All Institutions: Kandinsky Lab
The paper presents a practical and effective method for adding voice cloning to T2AV models with minimal architectural changes, achieving state-of-the-art speaker similarity at the cost of some transcription accuracy.
The paper proposes a method to add voice cloning capabilities to existing Text-to-Audio-Video (T2AV) diffusion models. The core innovation is architectural minimalism: adding a single zero-initialized linear layer to inject a global speaker embedding (via FiLM) and prepending reference audio latents to the input sequence. This allows the model to condition on a reference voice without retraining the entire backbone from scratch. The approach leverages the existing self-attention mechanisms for latent prepending and introduces a new modulation path for the speaker embedding. The methodology is sound and builds logically on existing practices in TTS (speaker embeddings) and diffusion (latent conditioning), but the combination within a T2AV context is a novel application. The zero-initialization ensures stability during the fine-tuning phase.
The evaluation is conducted on a benchmark of 674 speaker-text pairs across 30 speakers from the VCTK corpus. The authors compare their method against five strong baselines, including dedicated TTS models (Qwen3-TTS, XTTS-v2, IndexTTS2) and a joint audio-video model (NAVA). Results show that the proposed method achieves higher speaker-encoder cosine similarity (SECS) across three verification networks (ECAPA-TDNN, WavLM-SV, Resemblyzer) compared to baselines. However, this comes at the cost of higher Word Error Rates (WER), indicating a trade-off between speaker fidelity and transcription accuracy. The paper also demonstrates that the audio-only path can be run independently for a ~30x speed-up with minimal loss in speaker similarity. The "no regression" analysis shows that fine-tuning does not degrade the base model's performance on text-to-audio generation.
The paper provides detailed implementation details, including hyperparameters (learning rates, optimizer settings, diffusion steps), architecture dimensions, and training schedules. The use of open-source components (VCTK, Qwen3-TTS encoder, ECAPA-TDNN, etc.) aids reproducibility. However, the base T2AV model (Kandinsky 5.0) is described as having an "internal corpus" and the specific checkpoints are not publicly linked in the text provided, which may hinder exact replication. The code is not explicitly linked in the provided text.
The primary limitation is the trade-off between speaker similarity and speech intelligibility (WER). The model tends to preserve acoustic quirks of the reference, which can lead to hallucinations or errors in the generated text, especially on short or difficult prompts. The performance is also sensitive to reference length and language matching. The method relies on the availability of a pre-trained T2AV model, limiting its applicability to models that have not been trained on such data.
This work enables more personalized and controllable audio-visual generation, which has applications in content creation, dubbing, and virtual avatars. However, the ability to clone voices raises significant ethical concerns regarding misuse for deepfakes or non-consensual voice synthesis. The authors should have included a more robust discussion on these risks and potential mitigation strategies. The paper presents a practical and effective method for adding voice cloning to T2AV models with minimal architectural changes, achieving state-of-the-art speaker similarity at the cost of some transcription accuracy.
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller on an independent 2,000case text-audio dataset. MD-Judge supplies training rewards and adaptive search feedback, while a separate, non-adaptive Llama Guard 3 evaluator alone labels final outcomes. On 520 held-out AdvBench objectives, ARENA achieves FDR/PSR of 87.9/100.0%, 71.5/96.3%, 68.1/100.0%, and 75.4/98.5% on Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPTAudio, respectively. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.
Primary: Unknown
All Institutions: Unknown
[One sentence main contribution]. [The paper presents ARENA, a novel automated red-teaming framework that uses closed-loop feedback and preference optimization to generate effective audio-grounded jailbreaks for Large Audio-Language Models, demonstrating significant vulnerabilities in current safety mechanisms across multiple state-of-the-art models.]
The paper proposes ARENA, a closed-loop automated red-teaming framework for Large Audio-Language Models (LALMs). The core methodology involves training a controller LLM to generate text-safe queries paired with audio prompts (either speech or environmental sound) that induce harmful compliance in target models. The training utilizes a 2,000-case seed pool, employing Reward-Weighted Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) based on feedback from an MD-Judge model. During inference, the controller iteratively refines prompts based on judge feedback until a success threshold is met. A key technical distinction is the separation of the search feedback mechanism (MD-Judge) from the final evaluation metric (Llama Guard 3), which is intended to prevent overfitting to the judge. The approach addresses the specific challenge of "audio-grounded" jailbreaks where safety is conditional on the audio modality.
The authors evaluate ARENA on four LALMs: Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPT-Audio. They use 520 held-out AdvBench objectives. The results show high Fault Detection Rates (FDR), ranging from 68.1% to 87.9%, significantly outperforming static baselines like AJailBench and JALMBench. The paper includes ablation studies on refinement budgets, target sampling parameters, and audio variant counts, demonstrating that feedback-based refinement and audio variation substantially improve attack discovery. Transferability analysis shows that attacks found on one model can partially transfer to others. The evaluation is comprehensive, covering multiple models and providing detailed failure analysis.
The paper provides a GitHub link for code. The methodology describes the training data construction (2,000 seeds), the reward shaping formulas, and the DPO setup. However, the specific versions of the TTS models (Piper, TangoFlux) and the exact configuration of the MD-Judge and Llama Guard 3 evaluators are not fully detailed in the text provided, which may hinder exact reproduction. The separation of training and evaluation data is clearly stated, which is good practice.
The primary limitation is the reliance on LLM-based judges (MD-Judge and Llama Guard 3) for both training feedback and final evaluation. While the separation is intended to mitigate this, LLM judges are known to have biases and inconsistencies, particularly with audio-grounded reasoning. The paper does not provide human evaluation of the generated jailbreaks or the harmfulness of the responses, which is a significant gap for safety research. Additionally, the effectiveness is limited to the specific audio synthesis models used; different TTS or audio generation models might yield different results. The focus on automated red-teaming also means it may miss nuanced social engineering attacks that require more complex, multi-turn human interaction.
This work has significant implications for the safety of multimodal AI systems. By exposing vulnerabilities in LALMs, it helps developers identify and patch safety gaps before deployment. However, the release of such a powerful red-teaming tool also raises dual-use concerns, as the generated jailbreaks could be misused by malicious actors. The paper responsibly frames this as a safety auditing tool, but the availability of the code and methodology requires careful consideration of access controls. [One sentence main contribution]. [The paper presents ARENA, a novel automated red-teaming framework that uses closed-loop feedback and preference optimization to generate effective audio-grounded jailbreaks for Large Audio-Language Models, demonstrating significant vulnerabilities in current safety mechanisms across multiple state-of-the-art models.]
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.
Primary: Indian Institute of Science Education and Research Bhopal
All Institutions: Indian Institute of Science Education and Research Bhopal, Vellore Institute of Technology Bhopal
This paper presents a novel, geometric approach to test-time adaptation for audio-text models, demonstrating that affine corrections in latent space can effectively mitigate severe acoustic noise without gradients or source data.
The paper proposes PRISM, a training-free, source-free Test-Time Adaptation (TTA) framework for Audio-Text Foundation Models (ATMs). The core theoretical contribution is the "Affine Noise Hypothesis," which posits that severe acoustic noise induces a low-rank affine shift in the multimodal latent space. To correct this, PRISM employs three closed-form geometric operations: Orthogonal Procrustes Cross-modal Alignment (OPCA) to align manifolds, Class-Conditioned Variance Deflation (CCVD) to remove noise-dominant directions via Fisher Linear Discriminant Analysis, and Per-Class Residual Translation. These are compiled into a static projection matrix via Affine Bias Regression (ABR). The approach is mathematically grounded in linear algebra and manifold learning. While the geometric intuition is sound, the novelty lies primarily in the specific combination and application of these techniques to the ATM domain, rather than the introduction of fundamentally new mathematical primitives. The "Polyphonic Trap" analysis is a valuable diagnostic contribution, identifying a specific failure mode where high within-class variance in polyphonic sounds is mistaken for noise.
The evaluation is conducted on UrbanSound8K, ESC-50, and DCASE/TAU 2019 datasets with injected noise. The results show significant improvements over zero-shot baselines and other TTA methods like PCA++ and TDA. Notably, PRISM outperforms the oracle-assisted ContextDA baseline, which is a strong result given that ContextDA has access to privileged noise annotations. The paper provides detailed ablation studies and sensitivity analyses. However, the evaluation relies heavily on synthetic noise injection. While the SNR sweep is thorough, the generalization to real-world, non-stationary acoustic environments (beyond the TAU corpus) is less rigorously demonstrated. The comparison with gradient-based TTA methods highlights the speed advantage but does not fully explore the accuracy-latency trade-off in dynamic streaming scenarios where batch sizes might be smaller than the calibration buffer.
The paper provides detailed algorithmic steps, hyperparameters (K=60, p=0.8, etc.), and implementation details (LAION-CLAP checkpoint). The closed-form nature of the solution enhances reproducibility. However, the code is not publicly linked in the provided text, and the specific prompt templates used for text prototypes are only partially described ("20 diverse prompt templates"). The reliance on a specific foundation model (CLAP) limits direct generalizability to other architectures without adaptation.
The primary limitation is the "Polyphonic Trap," where the method fails for broadband, spectrally dense classes like street music. Although the authors propose Confidence-Aware Regression (CAR) to mitigate this, it adds complexity and the method still struggles with classes where semantic variance overlaps with noise subspace geometry. Additionally, the method assumes the noise distortion is low-rank and affine; if the acoustic environment induces complex, non-linear manifold warping that violates this hypothesis, performance may degrade. The method also requires a calibration batch, which may not be feasible in strictly real-time, single-sample inference scenarios without a warm-up period.
This work contributes to the robustness of audio foundation models in real-world, noisy environments, which is critical for applications like assistive listening, environmental monitoring, and mobile audio search. By providing a computationally efficient, training-free solution, it lowers the barrier for deploying robust ATMs on edge devices. The identification of the Polyphonic Trap offers insights into the limitations of subspace-based denoising for complex audio classes, guiding future research in geometric deep learning for audio. This paper presents a novel, geometric approach to test-time adaptation for audio-text models, demonstrating that affine corrections in latent space can effectively mitigate severe acoustic noise without gradients or source data.
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spatial capture capabilities, such as irregular or sparse arrays, as classical linear methods fail to reconstruct high-order spatial detail. We introduce a diffusion-based generative framework that models the statistical properties of HOA RIRs. This enables device-agnostic encoding from arbitrary microphone arrays, potentially unseen during data measurement. Our approach incorporates a posterior sampling procedure that enforces consistency between the estimated signals and the measurements while plausibly reconstructing spatial information that is unobservable from the limited measurements alone. Experiments on simulated data demonstrate that our method outperforms linear and neural baselines, achieving accurate HOA RIR estimation up to 12th order. A listening test with binaural renderings, including both simulated and measured RIRs, further confirms that the proposed method yields higher perceptual similarity to reference Ambisonics RIRs than all baselines. The flexibility and accuracy of the proposed framework opens new possibilities for scalable acoustics simulations.
Primary: Reality Labs Research, Meta
All Institutions: Reality Labs Research, Meta, Aalto University, Acoustics Lab
This paper presents a significant advancement in spatial audio processing by effectively leveraging diffusion models to solve the ill-posed problem of high-order Ambisonics encoding from sparse measurements, achieving state-of-the-art perceptual and objective performance.
The paper proposes a novel application of diffusion models to the ill-posed inverse problem of encoding Room Impulse Responses (RIRs) into High-Order Ambisonics (HOA) from sparse or irregular microphone arrays. The core methodological contribution is the integration of a device-agnostic diffusion prior with a posterior sampling procedure that enforces data consistency via a range-projected likelihood term. The use of a hybrid time-frequency/time-domain architecture (NCSN++ backbone with a U-Net refinement stage for early reflections) is a well-reasoned design choice given the distinct temporal characteristics of RIRs. The introduction of a compressed spectrogram distance metric for the likelihood guidance is a technically sound innovation to emphasize weak high-order components. The formulation separates the device-agnostic prior from the device-specific likelihood, which is a strong theoretical foundation for generalization.
The experimental evaluation is rigorous and comprehensive. The authors utilize two datasets: a large-scale internal FDTD simulation dataset and the public Treble-10 dataset. They compare against three strong baselines: Linear Least-Squares, a Time-Dependent Neural Encoder, and a Conditional Diffusion model specific to the device. The inclusion of both objective metrics (EDC, NPM) and a subjective listening test (MUSHRA-style with binaural rendering) provides robust validation. The results demonstrate clear superiority in both perceptual similarity and objective metrics, particularly in preserving early reflections and late reverberation tails. The ablation studies effectively isolate the contribution of the range-projection constraint and the hybrid architecture.
The paper provides sufficient detail for reproduction, including dataset descriptions, model architectures (NCSN++ based), training hyperparameters (AdamW, learning rate, batch size), and specific implementation details like the compressed spectrogram definition. The use of standard libraries (PyTorch) and well-known architectures aids reproducibility. However, the reliance on an internal FDTD dataset for the primary training results limits independent verification of the scale of the results, though the public Treble-10 results offer some ground truth.
The primary limitation is the distribution gap between simulated training data and real-world measured RIRs, which the authors acknowledge and demonstrate leads to performance drops on measured data (Eigenmike-64). The method assumes a highly accurate Array Transfer Function (ATF), which may not hold in real-world deployments with calibration errors. Additionally, the iterative sampling process is computationally intensive, making it less suitable for real-time applications without significant acceleration techniques. The evaluation is currently limited to a single device configuration (Aria Glasses) for the device-agnostic claim, though the method is theoretically general.
This work has significant implications for spatial audio processing, particularly for Virtual Reality (VR) and Augmented Reality (AR) applications where scalable and accurate acoustic simulation is crucial. By enabling high-quality HOA encoding from sparse arrays, it lowers the barrier for capturing spatial audio with wearable devices. It also facilitates the generation of large-scale training data for spatial audio models. The device-agnostic nature of the approach promotes interoperability across different hardware platforms. This paper presents a significant advancement in spatial audio processing by effectively leveraging diffusion models to solve the ill-posed problem of high-order Ambisonics encoding from sparse measurements, achieving state-of-the-art perceptual and objective performance.
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.
Primary: Universitat Pompeu Fabra (Music Technology Group)
All Institutions: Universitat Pompeu Fabra, Music Technology Group
This paper makes a significant contribution to the understanding of representation quality in music foundation models by introducing a systematic layer-wise analysis and a novel pitch-transposition equivariance metric, demonstrating that intrinsic properties can effectively guide layer selection, particularly for tonal tasks where standard metrics fail.
The paper presents a rigorous and systematic layer-wise analysis of 12 music foundation models across three distinct pre-training paradigms (masked, autoregressive, contrastive). The methodology is sound, employing a comprehensive suite of label-free intrinsic metrics (Intrinsic Dimension, Curvature, Anisotropy, Effective Rank, LiDAR, InfoNCE) to characterize representation geometry. The key methodological contribution is the introduction of Pitch-Transposition Equivariance (PTE), a novel metric designed to capture tonal structure that standard geometric metrics miss. The approach of correlating these intrinsic properties with downstream probe performance across 15 diverse MIR tasks provides a robust framework for understanding representation quality without relying solely on task-specific supervision.
The experimental evaluation is extensive, covering a wide range of downstream tasks including tonal, rhythmic, timbral, semantic, and similarity tasks. The results are well-supported, demonstrating that while standard metrics correlate well with non-tonal tasks, they fail for tonal tasks, necessitating the new PTE metric. The paper effectively demonstrates that intrinsic metrics can serve as effective proxies for layer selection, often outperforming trainable multi-layer fusion methods, particularly in low-data regimes. The analysis of depth-wise trends across different model families provides valuable insights into how representation properties evolve.
The paper provides significant detail on the models, datasets (MTG-Jamendo, GiantSteps, NSynth, etc.), and evaluation protocols. The code and extended results are available on the project page, enhancing reproducibility. The use of standard datasets and publicly available models facilitates independent verification.
The study is correlational; it identifies properties associated with good performance but does not establish causality. The metrics are evaluated on frozen representations, so their utility for guiding pre-training or fine-tuning is not directly addressed. The analysis is limited to 12 models, which, while diverse, may not cover all architectural variations. The PTE metric, while promising, is specific to tonal tasks and may not generalize to other musical attributes.
This work provides practical guidelines for selecting layers in music foundation models, potentially reducing the computational cost of model evaluation and deployment. It advances the theoretical understanding of representation quality in audio models, bridging the gap between geometric analysis and practical MIR performance. The findings are relevant to the broader community working with self-supervised audio representations. This paper makes a significant contribution to the understanding of representation quality in music foundation models by introducing a systematic layer-wise analysis and a novel pitch-transposition equivariance metric, demonstrating that intrinsic properties can effectively guide layer selection, particularly for tonal tasks where standard metrics fail.
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.
Primary: NVIDIA Corporation
All Institutions: NVIDIA Corporation
VoiceChat-TTS presents a robust, low-latency, and interruptible streaming TTS system that effectively bridges the gap between high-quality offline synthesis and the real-time demands of full-duplex interactive agents. The comprehensive evaluation, including detailed latency analysis and interruption handling benchmarks, demonstrates its practical utility and sets a strong baseline for future research in continuous speech synthesis.
The paper proposes VoiceChat-TTS, a continuous, streamable TTS model built upon the Audio Flamingo 3-Chat architecture. The core technical contributions involve adapting a streaming decoder for full-duplex interaction. Key modifications include: 1) A Character-Aware Subword Encoder to handle out-of-vocabulary subwords from LLMs by converting them to character sequences processed by a shallow Transformer. 2) An interruption mechanism using explicit control tokens to halt generation and transition to silence without resetting the KV cache. 3) Audio Prompt Conditioning using a 3-second reference audio to stabilize speaker identity at the start of generation, addressing the lack of context in streaming settings. 4) A Mixture of Gaussian Estimation Head (MoGH) for faster RVQ token decoding. The approach is pragmatic, focusing on engineering solutions (gated fusion, specific tokenization, silence modeling) to enable real-time, interruptible speech synthesis. While not fundamentally novel in terms of architecture (it extends existing streaming decoders), the specific integration of these components for a unified duplex TTS system is a valuable engineering contribution.
The evaluation is comprehensive and well-structured. The authors compare VoiceChat-TTS against strong offline (Chatterbox-TTS, Qwen3-TTS) and streaming (Audio Flamingo 3-Chat) baselines. Metrics include CER, WER, SECS (speaker similarity), and Squim-MOS. The results show that VoiceChat-TTS achieves competitive quality, significantly outperforming the base Audio Flamingo 3-Chat decoder in intelligibility and quality. The paper also includes a dedicated interruption evaluation using the Full-Duplex-Bench (FDB) subset, measuring Stop Latency and Leakage. The latency analysis is particularly strong, providing detailed breakdowns of acoustic-token ITL and codec decoding time on specific hardware (RTX A6000), demonstrating a 2.1x speedup over a comparable baseline. The use of both synthetic and real conversational data for training is well-justified.
The paper provides significant detail for reproduction. It specifies the model size (977M parameters), the base architecture (Gemma 3-based), the codec configuration (12.5 Hz, 31-codebook RVQ), and the training stages (pretraining on single-turn data, fine-tuning on multi-turn). The code is publicly available via NVIDIA NeMo Speech, and the model checkpoint is on Hugging Face. The training data sources are listed (LibriTTS, HiFiTTS, synthetic data, Fisher). The evaluation protocol is also clearly defined, including the use of Silero VAD for interruption timing detection. This high level of detail ensures that the work is reproducible.
The authors acknowledge several limitations. First, the model lacks user-audio conditioning, meaning it cannot dynamically adapt its prosody to the user's speech characteristics or interruptions in real-time (it relies on text tokens). Second, there is a lack of controlled component-wise ablations; while preliminary experiments suggested cumulative gains, a systematic ablation study is missing. Third, speaker similarity (SECS) degrades over longer sequences of continuous generation, particularly for unseen speakers, indicating a challenge in maintaining long-context speaker consistency. Finally, the evaluation of "silence" generation relies on ASR-based metrics, which may not fully capture the perceptual quality of silence or the naturalness of the transition into/out of silence.
VoiceChat-TTS contributes to the development of more natural and responsive human-computer interaction systems. By enabling low-latency, interruptible speech synthesis, it facilitates the deployment of always-on voice assistants and conversational agents that behave more like human interlocutors. This has positive implications for accessibility, customer service, and companion AI. However, the ease of generating high-quality, realistic speech also raises concerns about potential misuse in deepfakes or deceptive interactions, although the model's reliance on text input mitigates some of the risks associated with direct audio-to-audio generation. VoiceChat-TTS presents a robust, low-latency, and interruptible streaming TTS system that effectively bridges the gap between high-quality offline synthesis and the real-time demands of full-duplex interactive agents. The comprehensive evaluation, including detailed latency analysis and interruption handling benchmarks, demonstrates its practical utility and sets a strong baseline for future research in continuous speech synthesis.
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing. Existing Text-to-Audio (T2A) systems either reduce quoted speech to unintelligible vocal murmur or delegate it to a separate TTS model with post-hoc mixing, which forfeits control over when speech occurs and how it interacts with the scene. We present VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects. At the architecture level, chunk-wise causal factorization with independent per-chunk noise levels lets audio be emitted through sliding-window streaming inference with KV caching at variable target durations; to enable inference at arbitrary chunk granularities, we further pretrain the model with randomized chunk boundaries. At the preference level, multi-reward Negative-aware FineTuning (NFT) jointly optimizes semantic fidelity, linguistic accuracy, aesthetic quality, and temporal grounding At the data level, to supply the missing supervision for vocal content, we build VoxCorpus, a large-scale corpus whose captions quote the verbatim transcript of embedded speech with time intervals, and VoxBench, an interval-annotated benchmark with a temporal-grounding metric. Experiments on four benchmarks spanning general audio, speech, and unified vocalized audio validate the effectiveness and efficiency of VoxAudio. Our code and demos are available at https://voxaudio.github.io.
Primary: Zhejiang University
All Institutions: Zhejiang University
The paper presents a technically sound approach to streaming audio generation with speech, introducing valuable data and benchmarks. The adaptation of autoregressive flow matching with chunk-wise causal factorization is a meaningful contribution to efficient audio synthesis, and the multi-reward alignment strategy offers a robust path for preference optimization in continuous domains.
The paper proposes VoxAudio, a causal autoregressive flow matching model for vocalized audio synthesis. The core technical contribution lies in adapting flow matching to a streaming, autoregressive setting. Specifically, it employs chunk-wise causal factorization with independent per-chunk noise levels, allowing for sliding-window inference with KV caching. This addresses the train-test gap often seen in causal diffusion/flow models by randomizing chunk boundaries during pretraining. The methodology also includes a multi-reward Negative-aware FineTuning (NFT) stage to align the model with human preferences across semantic, linguistic, aesthetic, and temporal grounding dimensions. The architectural adaptations (causal convolutions, chunk-wise causal attention) are sound and necessary for the stated streaming objective. The integration of NFT for audio generation is a notable extension of recent RLHF techniques to continuous flow-matching domains.
The authors introduce VoxCorpus, a new dataset with verbatim speech annotations and time intervals, and VoxBench, a benchmark with a temporal-grounding metric. Experiments are conducted on four benchmarks. The results claim superiority in speech fidelity and timing precision over baselines. However, the provided text is truncated, preventing a full assessment of the quantitative results (e.g., FAD, UTMOS, WER scores) and ablation studies. The claim of "outperforming current vocalized audio generation baselines" is significant but requires rigorous verification against strong baselines like separate TTS+T2A pipelines or unified models like Dasheng-AudioGen. The introduction of a new benchmark is a strong positive for the community, provided the annotation quality and metric validity are high.
The paper provides a code and demo link. The methodology describes the use of pre-trained components (Universe Audio VAE, T5, Whisper, CLAP), which aids reproducibility. The specific details of the NFT reward weights and the exact implementation of the randomized chunk boundary training are crucial for reproduction and are likely detailed in the full text (which is partially truncated here). The use of standard open-source models for evaluation (Whisper, CLAP) ensures that the evaluation protocol is replicable.
The paper does not explicitly discuss the computational cost of the autoregressive flow matching compared to non-causal diffusion models, which is a significant factor for practical deployment. The reliance on multiple external models (Whisper for WER, CLAP for semantic reward) for NFT introduces potential biases and error propagation. The quality of the VoxCorpus dataset, particularly the accuracy of the "verbatim transcript" and time intervals, is critical and any errors there would limit the model's potential. The truncation of the text prevents assessment of failure cases or specific limitations in handling complex acoustic scenes.
VoxAudio addresses a significant gap in audio generation by enabling intelligible speech within environmental soundscapes, which has applications in podcasting, video dubbing, and immersive media. The release of VoxCorpus and VoxBench provides valuable resources for the community. However, the ability to generate realistic speech in arbitrary contexts raises concerns about misuse for deepfakes or misinformation, although the environmental context may mitigate some immediate risks compared to pure TTS. The paper presents a technically sound approach to streaming audio generation with speech, introducing valuable data and benchmarks. The adaptation of autoregressive flow matching with chunk-wise causal factorization is a meaningful contribution to efficient audio synthesis, and the multi-reward alignment strategy offers a robust path for preference optimization in continuous domains.
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.
Primary: VUI Labs Research
All Institutions: VUI Labs Research
Luna-TTS introduces a scalable, block-diffusion-based TTS framework that achieves state-of-the-art quality and latency by adapting pretrained LLMs through progressive architectural changes and RL post-training, effectively resolving the latency-quality trade-off inherent in previous diffusion TTS systems.
The paper proposes a novel architecture for Text-to-Speech (TTS) by adapting Large Language Models (LLMs) into Diffusion Language Models (dLLMs). Specifically, it introduces a "block-causal" attention mechanism that allows for streaming generation while retaining the parallel denoising benefits of diffusion models. The approach involves a progressive adaptation from causal to bidirectional and finally to block-causal attention on a pretrained AR text LLM. A key technical contribution is the application of Group Relative Policy Optimization (GRPO) directly over the realized denoising trajectory, optimizing for content correctness and speaker similarity. The use of a semantic-distilled RVQ tokenizer where the first codebook is anchored to linguistic content is a significant methodological detail that facilitates the diffusion process. The transition from fully parallel masked diffusion to block-autoregressive diffusion for latency reduction is a well-reasoned engineering solution to a known bottleneck in diffusion-based generation.
The evaluation is comprehensive, covering both objective metrics (CER, WER, SIM) and subjective/human-rated metrics for emotion and non-verbal vocalization (NVV) control. The paper claims state-of-the-art performance on Seed-TTS-Eval and CV3-Eval, outperforming both open-source and leading commercial systems. The inclusion of a "warmed serving protocol" to measure real-time factor (RTF) and first-block latency adds practical relevance to the experimental setup. The comparison against commercial baselines strengthens the claim of production-readiness. However, the specific numerical results are redacted in the provided text, preventing a precise verification of the magnitude of improvement, though the qualitative claims are strong.
The paper provides detailed descriptions of the architecture, including the tokenizer design, the masked diffusion formulation, and the training schedule. The mention of "0.6B backbone lineage" and "1 million hours of speech" provides scale context. However, the specific hyperparameters for the GRPO stage and the exact implementation details of the block-causal attention masking are likely critical for reproduction and may be subject to interpretation. The lack of publicly released code or weights (implied by "none" for URLs) significantly hinders immediate reproducibility for the broader community, although the technical report format suggests a high level of detail.
The primary limitation is the reliance on a specific RVQ tokenizer and the assumption that the semantic anchoring of the first codebook generalizes well across all languages and domains. The block-diffusion approach, while reducing latency, still requires multiple denoising steps per block, which may not be as fast as single-step AR decoding in extreme low-latency scenarios. Furthermore, the complexity of training a diffusion LLM with RL post-training is significantly higher than standard AR TTS training, potentially limiting accessibility. The paper does not extensively discuss the failure modes of the NVV control or the robustness of the emotion conditioning under adversarial text inputs.
This work represents a significant step towards making diffusion-based generative models viable for real-time, production-grade audio applications. By bridging the gap between the quality of diffusion models and the latency requirements of streaming TTS, it expands the toolkit available for developers building voice assistants, gaming NPCs, and accessibility tools. The emphasis on expressive control (emotion, NVVs) aligns with the growing demand for more natural and engaging human-computer interaction. However, the potential for misuse in deepfake generation remains a concern, necessitating robust safety measures in deployment. Luna-TTS introduces a scalable, block-diffusion-based TTS framework that achieves state-of-the-art quality and latency by adapting pretrained LLMs through progressive architectural changes and RL post-training, effectively resolving the latency-quality trade-off inherent in previous diffusion TTS systems.
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, and existing evaluators typically predict scores in a single forward pass without providing readable explanations. We introduce MUSECRITIC, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores. MUSECRITIC follows a two-stage training pipeline: a teacher model first provides high-quality critiques for supervised fine-tuning, after which the fine-tuned model generates its own critiques for reward learning, mitigating distribution shift between training and inference. On an in-domain test set of 200 SongEval songs, MUSECRITIC reduces macro-averaged mean squared error from 0.2875 to 0.2316 and improves macro-averaged LCC, SRCC, and Kendall's tau to 0.9068, 0.8838, and 0.7178, respectively. On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%. Moreover, using MUSECRITIC with GRPO improves Muse-0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. These results demonstrate that critique-conditioned reward modeling reduces scoring error and provides an effective optimization signal for song generation. The project repository is available at https://github.com/WuqnEl/MuseCritic.
Primary: Fudan University
All Institutions: Fudan University
MuseCritic introduces a critique-conditioned reward modeling framework for long-form song generation, demonstrating that natural-language aesthetic critiques serve as effective intermediate representations for improving score prediction accuracy and enabling successful reinforcement learning alignment.
The paper proposes MuseCritic, a semi-scalar reward model for long-form song generation that integrates natural-language aesthetic critiques as an intermediate representation before predicting continuous scores. The methodology involves a two-stage training pipeline: first, supervised fine-tuning (SFT) of an audio-language model backbone using critiques generated by a powerful external teacher (Gemini-3-Pro) conditioned on expert scores; second, reward learning where the model generates its own critiques (self-generated) to mitigate distribution shift, followed by training a reward head to predict scores conditioned on these self-generated critiques. This "critique-then-score" approach is conceptually sound and aligns with recent trends in LLM-as-a-Judge and reasoning-enhanced reward modeling. However, the novelty is somewhat limited by the reliance on a black-box teacher for data generation and the use of standard LoRA fine-tuning on a large foundation model, rather than proposing a fundamentally new architectural primitive for audio processing.
The experimental evaluation is comprehensive and rigorous. The authors evaluate on in-domain SongEval metrics (MSE, LCC, SRCC, KTAU) and out-of-domain Music Arena preference accuracy. MuseCritic consistently outperforms baselines, including a retrained SongEval (UTMOS) baseline and Gemini-3.1-Pro. Crucially, the ablation studies effectively isolate the contributions of the critique generation, self-generated critiques vs. offline critiques, and SFT initialization. The downstream reinforcement learning experiment using GRPO to optimize Muse-0.6B with MuseCritic as the reward model demonstrates practical utility, showing improvements across multiple aesthetic metrics. The use of deterministic decoding and fixed-prefix evaluation windows adds robustness to the downstream results.
The paper provides detailed implementation details, including hyperparameters, training configurations, and data splitting strategies. The code repository is linked. The use of a specific external teacher (Gemini-3-Pro) for data generation is a potential reproducibility hurdle for others without access to that specific model version, but the prompt templates are provided. The data splitting is clearly defined.
The primary limitation is the computational cost and latency introduced by the autoregressive critique generation step, which makes inference slower than direct score regression. Additionally, the method relies heavily on the quality of the external teacher for initial data generation, which may introduce biases or hallucinations if not properly verified (though the authors claim human verification). The evaluation is somewhat limited to Chinese and English vocal songs, and the generalizability to other genres or languages is not fully explored.
This work contributes to the field of AI-generated content evaluation and alignment. By providing a more interpretable and potentially more accurate reward model for song generation, it facilitates better optimization of generative models via reinforcement learning. This can lead to higher quality AI-generated music and more reliable evaluation benchmarks. The emphasis on interpretable critiques also aligns with the broader goal of making AI systems more transparent and trustworthy. MuseCritic introduces a critique-conditioned reward modeling framework for long-form song generation, demonstrating that natural-language aesthetic critiques serve as effective intermediate representations for improving score prediction accuracy and enabling successful reinforcement learning alignment.
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.
Primary: Unknown
All Institutions: Unknown
The paper presents a novel counterfactual evaluation framework that rigorously distinguishes between attribute occurrence and instruction-attributable control in text-to-music models, revealing that high agreement rates often reflect data priors rather than genuine controllability, thereby providing a more accurate and actionable benchmark for the field.
The paper introduces a rigorous counterfactual evaluation framework for text-to-music models, specifically targeting the distinction between "prompted attribute agreement" (mere occurrence) and "instruction-attributable control" (causal influence of the prompt). The core methodological contribution is the "matched neutral--A--B contrast family" design. By generating audio with a neutral prompt (omitting the target attribute) and two treatment prompts (specifying different target values) using a shared seed and frozen adapters, the authors isolate the effect of the instruction from the model's inherent output priors. This approach is statistically sound and addresses a critical flaw in current evaluation practices where high agreement rates may simply reflect the model's bias toward common attributes (e.g., 4/4 time). The metrics proposed—Treatment Agreement, Enhancement ($\Delta$), and Margin—are well-defined and provide a granular view of controllability.
The evaluation is applied to three prominent open systems: ACE-Step 1.5, Stable Audio 3 Medium, and LeVo2. The experiments cover global key and beat grouping. The results are compelling and counter-intuitive to standard benchmarks: while ACE-Step and Stable Audio 3 show strong key control, LeVo2 shows little attributable response. More importantly, for beat grouping, the paper demonstrates that high four-beat agreement in Stable Audio 3 is largely inherited from neutral outputs (0.97 neutral rate), and the explicit instruction actually *decreases* the probability of four-beat grouping (negative $\Delta$). This finding fundamentally changes the empirical conclusion about these models' capabilities. The use of off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels adds significant robustness to the claims.
The paper provides extensive details on the experimental setup, including frozen native-interface adapters, specific decoding settings, and seed handling. The authors release a code-and-data package with instantiated cases, prompts, configurations, and analysis code. The use of deterministic seeds and frozen model revisions ensures that the generation process is reproducible. The detailed appendix provides literal renderings of inputs and full provenance, making it highly reproducible.
The evaluation is limited to global key and beat grouping. It does not address local structural edits, modulation, or continuous controls. The neutral contrast necessarily changes the rendered carrier (e.g., omitting a field vs. filling it), which introduces a confound between the absence of the attribute and the change in prompt structure. The authors acknowledge this and use A/B swaps and off-attribute placebos to mitigate it, but a factorial design crossing carrier templates would be stronger. The evaluation is also limited to the specific public interfaces of the models; internal architectural changes that might improve control are not explored.
This work has significant implications for the development and evaluation of generative audio models. By establishing that "agreement" is not sufficient evidence of "control," it provides a new standard for benchmarking. This can guide model developers to focus on improving genuine instruction-following capabilities rather than just aligning with data priors. It also helps practitioners choose models based on their actual controllability profiles (e.g., good at rare targets vs. common targets). The framework is generalizable to other multimodal generation tasks where output priors are strong. The paper presents a novel counterfactual evaluation framework that rigorously distinguishes between attribute occurrence and instruction-attributable control in text-to-music models, revealing that high agreement rates often reflect data priors rather than genuine controllability, thereby providing a more accurate and actionable benchmark for the field.
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM-Gen, an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. MiDashengLM-Gen represents a first approach for general text-to-audio generation with one end-to-end trained model. Empirical evaluations demonstrate that MiDashengLM-Gen drastically improves speech intelligibility over existing unified models. On the Seed-TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text-to-Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed-audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi-research/midashenglm-gen and https://huggingface.co/mispeech/midashenglm-gen, and the demo page is available at https://xingws.github.io/midashenglm-gen-demo/.
Primary: Xiaomi Inc.
All Institutions: MiLM Plus, Xiaomi Inc., X-LANCE Lab, Shanghai Jiao Tong University
[One sentence main contribution]. MiDashengLM-Gen introduces an end-to-end LLM-based autoregressive framework with per-token flow matching for unified mixed-audio scene generation, achieving state-of-the-art speech intelligibility and competitive mixed-audio quality. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents a significant step forward in unified audio generation by effectively integrating large language models with continuous flow matching. The key insight regarding the DiT width constraint provides valuable theoretical and practical guidance for training such hybrid architectures. The substantial improvement in speech intelligibility addresses a major bottleneck in previous unified models, making the generated audio more usable for practical applications. While the approach builds on existing components (LLMs, Flow Matching, Tokenizers), the specific integration for variable-length mixed-audio generation is novel and well-validated.
The paper proposes MiDashengLM-Gen, an end-to-end framework for unified audio scene generation (speech, music, sound effects). The core architectural innovation is replacing the frozen text encoder and non-autoregressive diffusion backbone of its predecessor (Dasheng AudioGen) with a pre-trained Large Language Model (Qwen3-1.7B) coupled with per-token conditional flow matching. The method utilizes a structured multi-view captioning approach to decompose audio scenes into semantic views (global, transcript, SFX, music, env, speaker). A key technical contribution is the identification of a convergence prerequisite: the DiT decoder width must strictly exceed the audio latent dimensionality. The approach integrates an audio-text alignment stage to map high-dimensional audio latents into the LLM's token space. While the combination of LLMs and flow matching is not entirely new, applying it to *unified mixed-audio scene generation* with *per-token* flow matching for variable-length output is a distinct methodological shift from previous fixed-length or disjointed pipeline approaches.
The evaluation is comprehensive, covering single-type (AudioCaps, MusicCaps) and mixed-type (MECAT) generation, as well as speech intelligibility (Seed-TTS, multilingual WER/CER). The results show significant improvements in speech intelligibility (WER drop from 12.15% to 2.79% on Seed-TTS) compared to the previous Dasheng AudioGen model. The model maintains competitive performance on mixed-audio metrics (FAD, FD, KL) on MECAT, although it trails dedicated sound-effect models like TangoFlux on pure sound effects, which is acknowledged. The ablation studies effectively validate the necessity of the audio-text alignment stage and the DiT width constraint. The use of standard benchmarks (Seed-TTS, MECAT) ensures fair comparison.
The paper provides detailed implementation details, including model sizes (Qwen3-1.7B, DiT 16 layers), training hyperparameters (batch size, LR, epochs), and data sources. Code and checkpoints are made publicly available via GitHub and Hugging Face, significantly enhancing reproducibility. The description of the structured captioning and the specific flow matching objective is sufficiently detailed for replication.
The authors acknowledge several limitations: variable-length generation is bounded by training data distribution (1-20 seconds); speech intelligibility still trails dedicated TTS systems (especially in low-resource languages); and the model supports only coarse speaker-style control without explicit voice cloning or fine-grained temporal control. The trade-off between mixed-audio coordination and single-source acoustic fidelity (e.g., trailing TangoFlux on sound effects) is also a noted limitation.
This work advances the field of generative audio by demonstrating that LLM-based autoregressive frameworks can effectively handle complex, mixed-audio scene generation with high speech intelligibility. This has significant implications for immersive media, gaming, and film production where coherent multi-source audio is required. The open-source release contributes to the community's ability to build upon unified audio generation models. [One sentence main contribution]. MiDashengLM-Gen introduces an end-to-end LLM-based autoregressive framework with per-token flow matching for unified mixed-audio scene generation, achieving state-of-the-art speech intelligibility and competitive mixed-audio quality. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents a significant step forward in unified audio generation by effectively integrating large language models with continuous flow matching. The key insight regarding the DiT width constraint provides valuable theoretical and practical guidance for training such hybrid architectures. The substantial improvement in speech intelligibility addresses a major bottleneck in previous unified models, making the generated audio more usable for practical applications. While the approach builds on existing components (LLMs, Flow Matching, Tokenizers), the specific integration for variable-length mixed-audio generation is novel and well-validated.
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. We further introduce a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.75x speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality-latency trade-off for real-time SE.
Primary: National Taiwan University
All Institutions: Academia Sinica, National Taiwan University, Kore University of Enna, University of Palermo, NVIDIA
The paper presents RT-SEMamba, a causal Mamba-based speech enhancement model with progressive knowledge distillation, offering a competitive quality-latency trade-off for real-time applications.
The paper proposes RT-SEMamba, a fully causal speech enhancement model utilizing Time-Frequency Mamba (TF-Mamba) blocks. The core methodological contribution lies in adapting the selective state-space model (SSM) architecture for strict real-time streaming, specifically by ensuring causality in the temporal dimension while maintaining bidirectional frequency modeling. The authors introduce a progressive knowledge distillation (KD) strategy to compress an 8-layer teacher into a 1-layer student. This involves distilling both complex spectral outputs (magnitude, phase, complex spectrum) and intermediate feature representations. The methodology is technically sound and addresses a relevant gap in applying Mamba architectures to streaming audio, where memory efficiency is critical. However, the novelty is somewhat incremental; adapting existing SSM blocks for causality and applying standard KD techniques are well-established practices in the field, though their specific combination and tuning for SE are valuable.
The experiments are conducted on the VoiceBank-DEMAND dataset, a standard but relatively small benchmark for SE. The results show that the distilled 1-layer student achieves 3.18 PESQ, outperforming a naive 1-layer baseline (3.06) and approaching the 8-layer teacher (3.32). The paper provides a detailed analysis of the quality-latency trade-off, demonstrating that KD effectively shifts the Pareto frontier. The inclusion of hybrid Mamba-Transformer ablations adds depth to the architectural analysis. However, the evaluation is limited to objective metrics (PESQ, CSIG, etc.) on a single dataset. The lack of subjective listening tests or evaluation on larger, more diverse benchmarks (like DNS Challenge datasets) limits the generalizability of the claims. The comparison with prior works is fair, but the performance gains, while consistent, are modest in absolute terms.
The paper provides sufficient architectural details, including layer configurations, loss functions, and training hyperparameters (e.g., ramp-up steps). The code is promised to be released on GitHub, which enhances reproducibility. The dataset and evaluation protocol are standard, facilitating independent verification. The description of the streaming inference setup, including buffer management, is clear enough for implementation.
The primary limitation is the reliance on the VoiceBank-DEMAND dataset, which may not reflect performance on more challenging, real-world noise conditions or larger-scale training data. The model's performance on unseen, complex acoustic environments is not thoroughly tested. Additionally, the "progressive" nature of the KD is described but the specific scheduling or adaptive mechanisms are standard, limiting the perceived innovation in the distillation strategy itself. The paper does not discuss potential failure modes of the Mamba blocks in handling very long-term dependencies compared to Transformers in the streaming context.
This work contributes to the development of efficient, real-time audio processing systems, which are crucial for accessibility technologies (hearing aids), immersive communications (AR/VR), and edge computing applications. By demonstrating that Mamba-based models can achieve competitive performance with lower latency and memory footprint than Transformers, it encourages the exploration of alternative sequence modeling architectures in audio AI. The open-source release will further facilitate research in this direction. The paper presents RT-SEMamba, a causal Mamba-based speech enhancement model with progressive knowledge distillation, offering a competitive quality-latency trade-off for real-time applications.
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.
Primary: The Chinese University of Hong Kong, Shenzhen
All Institutions: The Chinese University of Hong Kong, Shenzhen, LIGHTSPEED, Independent Researcher
[One sentence main contribution]. The paper introduces Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and video responses by decoupling visual intent planning (VTP) and acoustic timing (speech units) to enable efficient training from heterogeneous data sources. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The work is a significant engineering achievement in integrating large language models with speech and video generation for interactive dialogue. Its primary value lies in the system-level design: the VTP interface and the shared acoustic-temporal interface allow for modular training and inference. While not introducing a new foundational model, it effectively demonstrates how to overcome data scarcity in omni-modal dialogue by leveraging structured intermediate representations. The "Prefix Streaming" technique is a notable contribution to the specific problem of autoregressive video generation stability. The paper is well-written, thoroughly evaluated, and addresses a timely and important problem in multimodal AI.
The paper proposes Ex-Omni-2D, a framework for omni-modal dialogue that generates text, speech, and video. The core methodological contribution is the "Visual Thought Plan" (VTP), a structured textual representation of visual intent (scene, emotion, motion) generated by the LLM, and a "native multi-codebook speech unit" interface. The architecture factors the problem: the LLM generates VTP and text; a Speech Generator produces acoustic units; and a Video Generator (based on Wan2.1) synthesizes video conditioned on the reference image, VTP, and frame-aligned speech units. A key technical detail is the "Prefix Streaming" mechanism for the video generator, which attempts to mitigate cumulative degradation in autoregressive video generation by carrying a clean latent from the previous chunk. The approach is a sophisticated integration of existing components (LLM, TTS, Video Diffusion) rather than a fundamental architectural breakthrough in any single modality. The factorization to avoid paired dialogue-video data is a pragmatic engineering solution to a data scarcity problem.
The evaluation covers audio quality (PQ, CU, SIM), video quality (SC, IQ, DD), and synchronization (Sync-C). The paper demonstrates that the Teacher model achieves high quality but is slow, while the Streaming Student offers a trade-off. The ablation studies on VTP and speech conditioning are valuable, showing that the VTP improves subject consistency and synchronization compared to a neutral plan. However, the video quality metrics (SC ~94%, IQ ~67%) are competitive but not state-of-the-art for dedicated video generators, which is expected given the constraints of dialogue-native generation. The audio quality is strong, leveraging recent TTS advances. The evaluation protocol is rigorous, using established benchmarks like VoiceBench and OmniCharacter.
The paper provides detailed descriptions of the architecture, training stages, and hyperparameters. The use of open-source backbones (Qwen3, Wan2.1) facilitates reproduction. The specific "Prefix Streaming" mechanism and data construction pipelines are described in sufficient detail for replication. The project page link suggests code or at least detailed visualizations may be available, though the code repository itself is not explicitly linked in the text provided.
The authors acknowledge several limitations: the speaker similarity is not perfect, the VTP is not independently sufficient for video control (requiring acoustic conditioning), and there is a measurable trade-off between language capability and visual planning in the shared LLM channel. The Streaming Student still has significant latency (first video chunk >3s) and is not real-time. The quality-efficiency trade-off means the fast model is significantly lower quality than the teacher.
This work contributes to the field of embodied AI and virtual humans by providing a framework for generating coherent, multi-modal responses. It addresses the challenge of generating visually present dialogue agents, which has applications in customer service, education, and entertainment. The factorization strategy offers a blueprint for training complex multi-modal systems where paired data is scarce. [One sentence main contribution]. The paper introduces Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and video responses by decoupling visual intent planning (VTP) and acoustic timing (speech units) to enable efficient training from heterogeneous data sources. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The work is a significant engineering achievement in integrating large language models with speech and video generation for interactive dialogue. Its primary value lies in the system-level design: the VTP interface and the shared acoustic-temporal interface allow for modular training and inference. While not introducing a new foundational model, it effectively demonstrates how to overcome data scarcity in omni-modal dialogue by leveraging structured intermediate representations. The "Prefix Streaming" technique is a notable contribution to the specific problem of autoregressive video generation stability. The paper is well-written, thoroughly evaluated, and addresses a timely and important problem in multimodal AI.
The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Our model develops an intrinsic self-awareness by learning to quantify the physical deficiencies of acoustic signals through targeted self-supervised tasks. This learned uncertainty is then operationalized via a novel Confidence-Fused Decoding mechanism, which provides both high-level instructions and frame-level attention modulation to the LLM decoder. Our experiments confirm the effectiveness of this approach. The model sets a new state-of-the-art on whispered speech with a 17% relative CER reduction on AISHELL6-Whisper. At the same time, it directly addresses the reliability trade-off, with hallucination rates dropping from over 25% to 4.5%.
Primary: Qwen Business Unit of Alibaba
All Institutions: Qwen Business Unit of Alibaba
The paper presents a novel framework for robust whispered speech recognition by integrating self-supervised uncertainty perception into an Audio-LLM, achieving state-of-the-art results and significantly reducing hallucinations.
The paper proposes a "Whisper-Aware LLM" framework that integrates an Uncertainty Perception Module (UPM) into an Audio-LLM (Qwen2-Audio). The core innovation lies in using self-supervised tasks—F0 contour prediction and masked spectrum reconstruction—to quantify signal uncertainty. This uncertainty is then operationalized via a "Confidence-Fused Decoding" mechanism that injects global instruction embeddings and frame-level attention biases into the LLM decoder. The methodology is technically sound and addresses a specific, well-defined problem (whispered speech recognition) with a plausible mechanism for handling signal ambiguity. The use of physics-informed auxiliary tasks (F0 absence) to drive uncertainty estimation is a novel angle compared to standard confidence calibration methods. However, the integration of uncertainty into the attention mechanism via additive bias is a relatively simple modification, and the "self-awareness" claim is somewhat marketing-heavy for what is essentially a multi-task learning setup with a specific decoding strategy.
The experimental evaluation is comprehensive, covering whispered speech (AISHELL6-Whisper, wTIMIT), general ASR (AISHELL-1, LibriSpeech), and a novel "Noise Hallucination Set" to evaluate reliability. The results show significant improvements in CER/WER on whispered speech and a drastic reduction in hallucination rates (from >25% to 4.5%). The ablation studies support the contribution of the UPM and the decoding mechanism. The inclusion of a hallucination metric is a strong point, addressing a critical gap in current ASR evaluations. The comparison against strong baselines (Whisper-v3, Qwen2-Audio, Seed-ASR) is appropriate. The 17% relative CER reduction is a substantial empirical gain.
The paper provides detailed implementation details, including the base model (Qwen2-Audio), UPM architecture (1D-CNN + Transformer), training stages, and hyperparameters (LoRA rank, learning rates). The dataset construction is described, though the "pure noise" collection is less defined. The three-stage training strategy is clearly outlined. However, the code is not released, and some details regarding the "Noise Hallucination Set" creation (source of non-speech sounds) are vague. Reproducibility is moderate to high, assuming access to the base Qwen2-Audio weights and the described datasets.
The paper does not extensively discuss the computational overhead of the UPM or the decoding mechanism during inference. The generalization of the UPM to other types of speech degradation (e.g., heavy noise, reverberation) beyond whispering is not thoroughly explored, although some general ASR results are provided. The reliance on F0 prediction for uncertainty might be less effective for non-tonal languages or speech with very low pitch, though the masked spectrum task helps mitigate this. The "hallucination" metric, while useful, is specific to the authors' curated dataset; broader standard metrics for hallucination in ASR are still evolving.
This work has significant implications for robust speech recognition systems, particularly in low-resource or noisy environments where whispered speech might occur (e.g., libraries, hospitals, late-night settings). By addressing the accuracy-reliability trade-off, it contributes to the development of more trustworthy AI systems. The methodology of using self-supervised uncertainty perception for decoding control could be adapted to other modalities or tasks requiring robustness to signal degradation. The paper presents a novel framework for robust whispered speech recognition by integrating self-supervised uncertainty perception into an Audio-LLM, achieving state-of-the-art results and significantly reducing hallucinations.
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle to achieve these two goals and mainly rely on supervised fine-tuning, yielding sub-optimal performance. While reinforcement learning (RL) shows promise, applying it to TDAC faces two main challenges: (1) existing rewards are too coarse to supervise multi-event, multi-attribute, and multi-relation descriptions in a fine-grained manner; and (2) temporal supervision is difficult for free-form captions, where flexible event-time expressions make reliable event-time correspondence challenging. To address these challenges, we propose AudioMap, a novel RL-based TDAC framework, which shifts to a unified cloze-and-choice reward paradigm. Specifically, we introduce the Evidence Sufficiency Reward (ESR) with an asymmetric hierarchical scoring mechanism to promote fine-grained accuracy and descriptive richness across diverse acoustic dimensions. Furthermore, we design the Event-Conditioned Temporal Reward (ECTR) to structurally bind timestamps to event semantics via temporal IoU, accompanied by a dual-curriculum learning strategy to facilitate the training process. Finally, to support this task, we construct the first time-aware fine-grained audio captioning dataset, AudioMapCap-44K, which contains 44K carefully annotated captions. Extensive experiments across diverse benchmarks show that AudioMap achieves state-of-the-art (SOTA) performance among open-source models and delivers competitive or superior results relative to proprietary models. Project page and release updates are available at https://github.com/ryysayhi/AudioMap.
Primary: Kling Team
All Institutions: Kling Team
AudioMap introduces a novel RL-based framework for time-aware dense audio captioning, utilizing a cloze-and-choice reward paradigm to enhance fine-grained semantic coverage and temporal grounding, achieving state-of-the-art results among open-source models.
The paper proposes AudioMap, an RL-based framework for Time-Aware Dense Audio Captioning (TDAC). The core methodological contribution is a shift from coarse scalar rewards to a "cloze-and-choice" paradigm. Specifically, it introduces the Evidence Sufficiency Reward (ESR), which uses a frozen examiner to generate multiple-choice questions about acoustic details, and the Event-Conditioned Temporal Reward (ECTR), which aligns event descriptions with timestamps using temporal IoU. The approach employs Group Relative Policy Optimization (GRPO) with a dual-curriculum strategy (input duration and reward complexity). The methodology is technically sound and addresses specific pain points in dense captioning (lack of fine-grained supervision and temporal grounding). However, the use of cloze tests for reward modeling is not entirely new in the broader LLM literature (e.g., Omni-Cloze), though its application to dense audio captioning with temporal constraints is a novel adaptation. The reliance on a large frozen examiner (Qwen3.6-27B) for reward computation is a significant computational overhead.
The authors construct a new dataset, AudioMapCap-44K, with 44K high-quality captions. Experiments show SOTA performance among open-source models on Omni-Cloze, MMSU, MMAU, and TACOS benchmarks. The results are competitive with proprietary models like Gemini-3.1-Pro. The ablation studies effectively demonstrate the contribution of ESR (semantic coverage) and ECTR (temporal grounding). The user study provides subjective validation. The evaluation is comprehensive, covering semantic, QA-based, and temporal metrics. The comparison with proprietary models is strong.
The paper provides a GitHub link. The methodology describes the training setup (Qwen2.5-Omni initialization, GRPO details, curriculum stages). However, the reliance on proprietary models for dataset construction (Gemini-3.1-Pro, GPT-4.1) and the specific prompt engineering for the cloze question generation might make exact replication difficult without the specific prompts or access to those models. The code release is a positive sign.
The method is computationally expensive due to the need for a large examiner model during RL training. The dataset construction relies heavily on proprietary LLMs, which may introduce biases or limitations in the diversity of the data. The temporal grounding is still dependent on the examiner's ability to extract timestamps, which can be error-prone for concurrent events. The paper does not extensively discuss the failure cases of the temporal reward in complex acoustic scenes.
This work advances the field of audio understanding by enabling more detailed and temporally precise audio descriptions, which is crucial for accessibility, multimedia retrieval, and embodied AI. The dataset AudioMapCap-44K is a valuable resource for the community. The RL-based approach offers a generalizable framework for other dense captioning tasks. AudioMap introduces a novel RL-based framework for time-aware dense audio captioning, utilizing a cloze-and-choice reward paradigm to enhance fine-grained semantic coverage and temporal grounding, achieving state-of-the-art results among open-source models.
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduce MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation. MusicLayout describes a musical piece as a time-aligned layout of sections, textures, repetitions, variations, and instrument-level arrangements, serving as an interpretable planning layer between textual intent and the generated music. We integrate MusicLayout into a text-to-music framework built on a unified autoregressive formulation, where the model first generates a MusicLayout representation and subsequently predicts audio tokens conditioned on this representation within a single sequence. The resulting MusicLayout can be inspected and modified prior to audio generation, providing a mechanism for layout-level structural control. We evaluate MusicLayout through layout-conditioned generation, layout manipulation experiments, and matched-data ablations, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control.
Primary: Zhejiang University
All Institutions: Zhejiang University, The Chinese University of Hong Kong, Shandong University, Ant Group
MusicLayout introduces an explicit, time-aligned intermediate representation for controlling musical structure in text-to-music generation, enabling interpretable planning and layout-level control within a unified autoregressive framework.
The paper proposes MusicLayout, an explicit intermediate representation for text-to-music generation. The core methodological contribution is the design of a structured, time-aligned token sequence that describes musical sections, textures, variations, and instrument arrangements. This layout is generated by a unified autoregressive language model (based on ACE-Step 1.5) before audio tokens are predicted. The approach effectively introduces a "planning" layer, analogous to spatial layout planning in image generation, but adapted for the temporal domain of music. The serialization grammar is well-defined, and the integration into the autoregressive sequence is straightforward. However, the novelty is somewhat incremental; it applies a known paradigm (explicit intermediate representation) to a new modality (music) without introducing fundamentally new algorithmic mechanisms for the generation process itself. The reliance on MIDI-synthesized audio for training the layout structure is a significant methodological constraint, as it decouples the structural planning from the acoustic fidelity of the final output.
The evaluation is comprehensive, covering objective metrics (FAD, CLAPScore, PaSST-KL, SSIM, SCM Energy Distance, boundary agreement) and subjective human evaluation. The authors include strong baselines (MusicGen, ACE-Step, Stable Audio) and, crucially, matched-data controls (finetuned ACE-Step without layout, shuffled layout training/inference) to isolate the effect of the explicit planning. The results show that MusicLayout improves structural organization (boundary agreement) and can be manipulated to control structure. However, the acoustic quality metrics (FAD, CLAP) are generally inferior to or comparable with baselines, which the authors attribute to the use of MIDI-synthesized training data. The evaluation on MuChin (real audio) shows a performance drop, highlighting the domain gap. The subjective evaluation supports the claim of improved structural control and text consistency for experienced listeners. The ablation studies are rigorous and help validate the specific contribution of the layout tokens.
The paper provides detailed implementation details, including the model architecture (ACE-Step 1.5 backbone), training settings (batch size, learning rate, GPU count), and the specific datasets used (FreeMIDI, MidiCaps, MuChin). The annotation pipeline for creating MusicLayout from MIDI is described in an algorithmic format. The code for the layout extraction and the model adaptation is not explicitly linked, but the description is sufficient for a competent researcher to reproduce the method. The use of open-source components (FluidSynth, PrettyMIDI, ACE-Step) aids reproducibility.
The primary limitation is the reliance on MIDI-synthesized audio for training the structural planner. This creates a domain gap when evaluating on real audio (MuChin), where performance degrades. The system cannot edit existing audio or regenerate selected regions; it only allows pre-synthesis layout manipulation. The vocabulary for sections, textures, and instruments is closed and may not cover all musical styles or nuances. The quality of the generated layout depends on the text prompt, and if the prompt is ambiguous, the generated layout may not align with the user's intent. The acoustic fidelity is limited by the quality of the MIDI synthesis used in training, although the DiT renderer is frozen and presumably capable of higher quality if trained on better data.
This work contributes to the field of controllable generative AI by demonstrating the value of explicit intermediate representations for complex, structured data like music. It opens up new avenues for human-in-the-loop music creation, where users can plan the structure of a piece before generation. This could lower the barrier to entry for music production and enable more precise artistic control. However, the reliance on MIDI data and the potential for generating low-quality audio from poor MIDI synthesis are practical concerns. The work also raises questions about the generalizability of such explicit planning methods to other domains. MusicLayout introduces an explicit, time-aligned intermediate representation for controlling musical structure in text-to-music generation, enabling interpretable planning and layout-level control within a unified autoregressive framework.
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Primary: ServiceNow
All Institutions: ServiceNow
This paper introduces a linguistically grounded, dimension-level meta-evaluation benchmark for TTS, revealing significant gaps in the diagnostic capabilities of current MOS predictors and Audio-LLM judges, thereby providing a crucial foundation for more interpretable and targeted speech quality assessment.
The paper proposes a novel meta-evaluation framework for Text-to-Speech (TTS) systems, moving beyond holistic Mean Opinion Score (MOS) predictors. The core methodological contribution is a linguistically grounded annotation schema that deconstructs "naturalness" into 10 distinct perceptual dimensions across word, prosodic, and paralinguistic levels. The authors construct a dataset of 860 utterances with controlled errors introduced via LLM-altered IPA, acoustic manipulation, and API-enforced emotion tags. They then benchmark four neural MOS predictors and four Audio-LLM judges against this human-annotated ground truth. The methodology is rigorous in its construction of a diagnostic benchmark, though the error generation methods (particularly acoustic manipulation) represent upper-bound cases that may not reflect natural system failures. The approach to evaluating Audio-LLMs via varied prompting conditions (schema-guided vs. isolated) is a significant methodological step for understanding model capabilities.
The experimental design is comprehensive, covering a wide range of current state-of-the-art evaluators. The results clearly demonstrate that MOS predictors collapse onto acoustic signal quality rather than linguistic naturalness, and that Audio-LLM judges exhibit selective, prompt-dependent detection capabilities. The use of Krippendorff's alpha to assess inter-annotator reliability adds robustness to the human ground truth. However, the dataset size (860 samples) is relatively small for such a high-dimensional evaluation space, and the reliance on synthetic error injection limits the generalizability of the findings to real-world TTS system outputs. The correlation analysis between automated scores and human labels is well-executed but reveals a significant gap in current technology's ability to diagnose specific linguistic errors.
The paper states that the dataset, annotation schema, and evaluation code are publicly released, which strongly supports reproducibility. The detailed description of the error generation pipeline, including specific tools (Praat, Parselmouth) and model versions (Cartesia Sonic-3, GPT-5), allows for replication of the dataset construction. The clear definition of the 10 dimensions and the binary rating scale facilitates consistent annotation by future researchers.
The primary limitation is the reliance on synthetic errors. Acoustic manipulations (e.g., F0 contour changes, duration scaling) create extreme cases that may not be representative of the subtle, complex errors found in production TTS systems. Additionally, the dataset is limited to English, restricting the generalizability of the schema to other languages with different phonological and prosodic structures. The small sample size per dimension (ranging from 46 to 128) may limit the statistical power of the findings for less common error types. The evaluation of Audio-LLMs is also constrained by the specific prompts tested, and the "black box" nature of these models means the exact reasoning behind their scores remains opaque.
This paper has significant implications for the development and deployment of TTS systems. By providing a diagnostic benchmark, it enables developers to identify specific weaknesses in their models rather than relying on opaque holistic scores. This can lead to more targeted improvements in TTS technology. The release of the dataset and schema also contributes to the broader ML community by providing a standardized resource for evaluating speech quality along linguistically grounded dimensions. This work highlights the current limitations of automated evaluators and underscores the need for more sophisticated models that can understand and diagnose linguistic nuances in speech. This paper introduces a linguistically grounded, dimension-level meta-evaluation benchmark for TTS, revealing significant gaps in the diagnostic capabilities of current MOS predictors and Audio-LLM judges, thereby providing a crucial foundation for more interpretable and targeted speech quality assessment.
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper studies cross-instrument MIDI velocity estimation in this label-scarce setting. Starting from a piano-trained velocity estimator, we recast target-instrument adaptation as predicting renderer-conditioned velocities whose rendering matches the dynamics of the performance audio. This adaptation can be driven by either differentiable synthesizers (Diff-Synth) or our proposed differentiable SoundFont proxies (Diff-SFProxy). We highlight the Diff-SFProxy: it supervises velocity through note-wise, loudness-related acoustic parameters rather than waveform reconstruction, focusing gradients on velocity-dependent behavior. Experiments on piano and guitar show that Diff-SFProxy is effective for cross-instrument MIDI velocity estimation, while waveform-domain Diff-Synth degrades performance.
Primary: University of Western Australia
All Institutions: University of New South Wales, University of Western Australia
The paper presents a robust and innovative solution to cross-instrument MIDI velocity estimation by introducing a differentiable SoundFont proxy that leverages perceptual loudness parameters for gradient-based adaptation, effectively overcoming the limitations of waveform-based methods in label-scarce settings.
The paper proposes a novel framework for cross-instrument MIDI velocity estimation in label-scarce settings. The core innovation is the "Diff-SFProxy," a differentiable proxy that maps MIDI note events to loudness-related acoustic parameters (Pitch-conditioned Harmonic Energy and Onset-window Spectral Flux) derived from SoundFont renders. This allows gradient-based adaptation of a piano-trained velocity estimator (VeloEst) to target instruments (guitar) without requiring ground-truth velocity labels. The methodology effectively addresses the "gradient mismatch" problem inherent in waveform-based differentiable synthesis (Diff-Synth), where timbral and room-acoustic mismatches dominate the loss landscape. The use of perceptual loudness metrics (Bark-scale specific/total loudness) for evaluation is a strong methodological choice, aligning the objective with the perceptual goal of velocity estimation.
The experimental design is rigorous and well-structured. The authors include a critical "velocity-recovery diagnostic" on synthetic audio to verify that the Diff-SFProxy gradients point in the correct direction, which is essential for trusting the adaptation process. The evaluation on real-world datasets (MAESTRO, SMD, GAPS, FL) demonstrates that Diff-SFProxy significantly outperforms both flat-velocity baselines and zero-shot transfer. Crucially, it shows that Diff-Synth often degrades performance on guitar, validating the paper's central hypothesis that parameter-space supervision is superior to waveform-space supervision for this specific task. The ablation studies on segment length and loss components further strengthen the findings.
The paper provides substantial implementation details, including FFT settings, optimizer hyperparameters, and architecture specifications for the Transformer proxy. The code is explicitly linked via GitHub. The use of standard datasets (MAESTRO, GAPS) and public SoundFonts enhances reproducibility. The description of the training protocol, including the mixture sampler and caching strategy, is clear enough for replication.
The approach relies on the availability of a suitable SoundFont for the target instrument, which requires manual selection or future automation. The recovered velocities are "renderer-conditioned," meaning they are optimal for the specific SoundFont used, not necessarily the canonical physical velocity of the performance. The method is currently limited to instruments with available SoundFonts and may struggle with complex polyphony or instruments with significant inharmonicity not captured by the proxy. The evaluation is limited to piano and guitar; generalization to other instruments (violin, wind) is speculative.
This work addresses a significant bottleneck in symbolic music processing: the lack of expressive velocity data for non-piano instruments. By enabling label-scarce adaptation, it facilitates the creation of more expressive MIDI representations for music generation, analysis, and education. The Diff-SFProxy framework could be extended to other expressive parameters (e.g., pedal, bow pressure), potentially enriching the symbolic music domain. The paper presents a robust and innovative solution to cross-instrument MIDI velocity estimation by introducing a differentiable SoundFont proxy that leverages perceptual loudness parameters for gradient-based adaptation, effectively overcoming the limitations of waveform-based methods in label-scarce settings.
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.
Primary: Unknown
All Institutions: Unknown
The paper introduces MADBench, a novel benchmark for component-aware audio deepfake detection that disentangles speech and environmental audio manipulation, providing rigorous evaluation protocols and insightful findings on detector behavior and cross-component interference.
The paper proposes MADBench, a novel benchmark for audio deepfake detection that disentangles speech and environmental audio components. The methodology involves a rigorous pipeline: source video separation using MossFormer2, generation of fake speech via TTS/VC models (F5-TTS, SeedVC, etc.), and generation of fake environmental audio via Text-to-Audio and Video-to-Audio models (AudioLDM2, MMAudio, etc.). A key methodological contribution is the introduction of a "scene-consistency" axis (matched vs. mismatched environmental audio) and the evaluation of cross-component interference. The approach is technically sound, addressing a genuine gap in existing benchmarks which often conflate audio modalities or focus solely on speech. However, the novelty is somewhat tempered by the fact that the core components (separation, TTS, TTA) are established techniques; the innovation lies primarily in the specific combination and the evaluation protocol rather than a new algorithmic breakthrough.
The experimental evaluation is comprehensive. The authors benchmark a wide range of models: pretrained AV detectors (AVH-Align, etc.), frozen AV encoders (ImageBind, PE-AV, CAV-MAE), and zero-shot Omni models (Qwen2.5-Omni, MiniCPM-o). They evaluate across multiple protocols (Binary, 4-way, Component-level, Scene-consistency). The results are insightful: they find that environmental audio manipulation is easier to detect than speech, and that fake environmental audio can obscure speech detection (cross-component interference). They also find that frozen AV encoders perform surprisingly well, often better than task-specific detectors, and that video input does not consistently improve manipulation detection but helps with scene consistency. The analysis is thorough, including ablation studies on input modality and scene consistency. The statistical reporting (AUC, EER) is appropriate.
The paper provides detailed descriptions of the dataset construction, including the source dataset (AVSpeech), the separation model (MossFormer2), the specific generation models used, and the quality control steps. The split strategy (7:1:2) and the balancing of generation settings are described. However, the paper does not provide a public link to the dataset or code in the text provided (URLs are "none"). The use of proprietary or hard-to-access models (e.g., specific checkpoints of F5-TTS, SeedVC, etc.) might pose reproducibility challenges for the exact generation pipeline, although the models themselves are generally available. The detailed protocol allows for reasonable reproducibility of the evaluation framework.
The paper acknowledges that the visual stream is held fixed, which limits the study to audio manipulation detection in authentic video. It does not address scenarios where both audio and video are manipulated. The dataset size (1,892 clips) is relatively small for deep learning benchmarks, which might limit the generalizability of the findings, especially for complex multimodal models. The reliance on automatic scene classification (Qwen2.5-VL) for taxonomy construction introduces potential noise, although manual verification was performed. The performance of zero-shot models is limited, which is expected but highlights the current gap in multimodal reasoning for forensic tasks.
This paper has significant broader impact. By providing a benchmark that reflects a realistic attack scenario (independent manipulation of speech and background), it enables more robust development of deepfake detection systems. The findings that environmental audio manipulation is easier to detect and can interfere with speech detection have practical implications for security and forensics. The benchmark will likely spur research into more robust, component-aware detection models. It also highlights the limitations of current zero-shot multimodal models in forensic reasoning, guiding future research directions. The paper introduces MADBench, a novel benchmark for component-aware audio deepfake detection that disentangles speech and environmental audio manipulation, providing rigorous evaluation protocols and insightful findings on detector behavior and cross-component interference.