Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: HKGAI
All Institutions: HKGAI
YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The paper proposes a unified architecture (YuE2) that bridges symbolic and audio music generation using an AR-NAR Mixture-of-Transformers (MoT). The core methodological contribution is the "symbolic planning" stage, where the model first generates a readable score (melody, harmony, form) before expanding into semantic tokens and finally acoustic latents via flow matching. This hierarchical approach allows for explicit control over composition. Additionally, the paper introduces two auxiliary models: MERT2 for semantic representation learning and SheetSage2 for symbolic supervision (lead-sheet transcription), which are used to train the main model on unaligned audio data. The integration of these components into a single checkpoint that supports editing and cover generation is a significant architectural advancement.
The evaluation is extensive, utilizing both objective benchmarks (WildSongBench, SongBench, MARBLE) and subjective expert listening tests. The results show YuE2 outperforming public baselines and being competitive with proprietary systems like Suno v4.5 and v5. The ablation study comparing generation with and without symbolic planning demonstrates a clear preference for the planned approach (49.3% vs 34.6%). The introduction of MERT2 and SheetSage2 is supported by state-of-the-art results on their respective benchmarks (MARBLE and transcription metrics). The inclusion of zero-shot cover generation and agentic editing case studies further validates the utility of the unified framework.
As a technical report from an industry lab (HKGAI), the paper likely lacks full open-source code and model weights at the time of release, which limits immediate reproducibility. However, the detailed description of the architecture (AR-NAR MoT, flow matching) and the specific benchmarks used provides a clear roadmap for replication. The reliance on proprietary datasets for training the auxiliary models may pose a barrier for independent researchers.
The paper relies heavily on expert listening tests, which can be subjective and difficult to scale. The comparison with proprietary systems (Suno) is limited to specific versions and may not reflect the full capability of those closed-source models. The "best-of-8" selection strategy, while effective for benchmarking, introduces a computational overhead that may not be practical for real-time applications. Additionally, the paper does not extensively discuss the computational cost of the multi-stage generation process compared to direct audio generation.
This work has significant implications for the music production industry by providing a tool that allows for both high-quality audio generation and explicit compositional control. The ability to edit scores and generate covers zero-shot opens up new possibilities for music creation, education, and collaboration. The unification of symbolic and audio domains could lead to more interpretable and controllable generative music systems, potentially influencing how AI is integrated into professional music workflows. YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: MIT CSAIL
All Institutions: MIT CSAIL, MIT RLE, KAIST GSCT
The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
The paper introduces a differentiable, GPU-accelerated acoustic simulator for the vocal tract based on the linearized Euler equations. The core technical contribution is the shift from time-domain finite difference methods to a frequency-domain formulation using transmission line matrices (TLM). This allows for parallelization across frequencies, resulting in a 70x speedup and, crucially, smoother gradient landscapes that facilitate stable optimization. The authors also propose a differentiable turbulence model using softplus gating and softmax-based constriction localization to handle consonants, unifying vowel and consonant synthesis. Finally, they leverage neural fields (implicit neural representations) to parameterize the vocal tract geometry, which acts as a regularizer to escape local minima in the non-convex inverse problem.
The experiments are extensive and well-designed. The authors validate the simulator's accuracy by comparing Frequency Domain Synthesis (FDS) against Time Domain Synthesis (TDS), showing significant improvements in SI-SDR, STOI, and PESQ metrics. They demonstrate the utility of the simulator in two novel applications: (1) a self-supervised autoencoder that maps audio to vocal tract area functions across 11 languages, outperforming paired-data baselines in intelligibility and speaker identity preservation; and (2) a generative model that reconstructs MRI videos of the vocal tract from speech alone, without paired data. The emergence of the IPA vowel chart from the latent space of the MRI generator provides strong evidence for the physical grounding of the model.
The paper provides detailed derivations of the physics, including the linearized Euler equations, Webster's equation, and the circuit interpretation. Specific hyperparameters, such as window lengths, hop sizes, and Reynolds number thresholds, are provided. The use of standard architectures like Wav2Vec 2.0 and StyleGAN2 aids reproducibility. However, the specific implementation of the differentiable turbulence model and the exact neural field architectures (RFF vs MFN) would benefit from more code-level details or a public repository link beyond the project page.
The model relies on a 1D acoustic tube approximation, which may not capture all 3D effects of the vocal tract, particularly for complex articulations or nasal sounds (though the nasal tract is noted as a limitation due to MRI data availability). The inverse problem is inherently ill-posed, meaning multiple geometries can produce the same sound, leading to potential ambiguities in reconstruction. The MRI reconstruction quality is limited by the resolution and coverage of the training MRI dataset.
This work has significant implications for linguistics, speech pathology, and medical imaging. By providing a physically grounded link between speech and articulatory geometry, it enables new tools for voice coaching, language acquisition studies, and non-invasive visualization of vocal tract movements. The self-supervised nature of the autoencoder and MRI generator makes these tools scalable and accessible without the need for expensive paired data collection. The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Primary: The Chinese University of Hong Kong
All Institutions: The Chinese University of Hong Kong, Microsoft Corporation
The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
The paper proposes DuraS2ST, a framework that reformulates duration-aligned speech-to-speech translation (S2ST) as a reasoning problem rather than an acoustic post-processing task. The core methodological contribution is the integration of Chain-of-Thought (CoT) reasoning into the speech generation pipeline. The model first generates an explicit textual rationale planning the target wording and phonetic length, followed by the synthesis of interleaved text-acoustic tokens. This is supported by the construction of DuraSet-440K, a large-scale corpus specifically designed for this purpose. The training paradigm is two-phase: Supervised Fine-Tuning (SFT) on the new corpus, followed by Group Relative Policy Optimization (GRPO). The RL phase introduces two specific technical innovations: the Duration Margin Reward (DMR), which acts as a soft constraint to balance translation quality with duration consistency, and Modality-Aware Reward Attribution (MARA), which prevents cross-modality reward contamination by assigning duration rewards only to acoustic tokens and quality rewards to both text and acoustic spans. This approach is technically sound and addresses a specific gap in current S2ST systems where duration control is often handled via black-box speed tokens or post-hoc time-stretching.
The experiments are conducted on the CVSS-T benchmark, which is disjoint from the training data, providing a valid out-of-domain evaluation. The paper compares DuraS2ST against strong baselines, including commercial models (GPT-4o, Qwen2.5-Omni, Kimi-Audio) and open-source SLMs (Step-Audio-2-mini). The results demonstrate that DuraS2ST achieves a superior balance between translation quality (BLEU, COMET) and duration consistency (SLC-0.2, SLC-0.4, MADE, MRDE). Notably, the method shows substantial gains in a zero-shot RL setting, highlighting the effectiveness of the reward design. The inclusion of both open-source and commercial baselines strengthens the experimental validation.
The paper provides a project page with a GitHub repository link. It details the training setup, including hyperparameters for SFT and GRPO, and describes the construction pipeline for DuraSet-440K in detail. The use of standard metrics (BLEU, COMET, SECS, SLC) and a public benchmark (CVSS-T) enhances reproducibility. However, the specific implementation details of the "duration-controllable TTS model" used for data synthesis are not fully specified in the main text, though likely detailed in the appendix or code.
The method relies on a specific base model (Step-Audio-2-mini-Think), and its generalizability to other SLM architectures is not extensively tested. The construction of DuraSet-440K is resource-intensive, requiring LLM-based translation candidate generation and duration-controllable TTS synthesis. The evaluation is limited to English-Chinese pairs; performance on other language pairs with different phonetic structures remains unknown. Additionally, the "zero-shot RL" claim needs careful interpretation as it likely refers to the RL phase without additional SFT on the target domain, but the model is still SFT'd on DuraSet-440K.
This work has significant implications for video dubbing, simultaneous interpretation, and real-time communication applications where audio-visual synchronization is critical. By treating duration alignment as a reasoning task, it offers a more interpretable and controllable approach compared to traditional acoustic constraints. The release of DuraSet-440K and the code will likely accelerate research in duration-controlled speech generation and multimodal reasoning. The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ฯ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io
Primary: The University of Texas at Austin
All Institutions: The University of Texas at Austin
The paper introduces a label-free, linear-predictive method for pruning audio tokens in LALMs before the language model runs, significantly improving efficiency and context window capacity. The technical contribution is strong, combining a novel insight about attention predictability with a practical, two-stage pruning algorithm that outperforms existing baselines across multiple models and tasks.
The paper proposes "Triage," a two-stage audio token pruning method for Large Audio Language Models (LALMs). The core innovation is the discovery that the attention an audio token will receive across the entire language model (LM) is linearly predictable from the encoder output alone, before the LM runs. The authors fit a simple linear map (the "prior") in closed form using ridge regression on unlabelled data. This prior is used in Stage 1 to prune tokens before the LM prefill, significantly reducing context window usage. Stage 2, applied to multiple-choice tasks, refines the ranking using attention observed at layer 2, combining the prior's prediction with observed attention via a "precision fusion" mechanism. The method is label-free for calibration, using self-consistency with the model's own full-audio outputs to set compression budgets. This approach is distinct from existing methods that rely on acoustic energy, position, or mid-prefill attention, which the paper demonstrates are weak predictors of final attention for audio tokens.
The experimental evaluation is extensive and rigorous. The authors test the method on 13 different LALMs from 5 families, demonstrating the generality of the linear prior (achieving Spearman correlation >= 0.69 on 11/13 models). They evaluate on both transcription (LibriSpeech, FLEURS, TEDLIUM) and multiple-choice (MMSU, DREAM, AudioMarathon-RACE) benchmarks. The results show that Triage outperforms strong baselines like DART, FastV, and HeadRouter, particularly at aggressive compression ratios. A key strength is the efficiency analysis, showing that Triage allows 4x more concurrent streams on a single GPU and extends the effective context window from ~22 to ~62 minutes for Qwen2.5-Omni-3B. The ablation studies convincingly show that both the prior and the stage-2 refinement contribute to performance.
The paper provides high reproducibility. The method relies on a closed-form linear fit, which is computationally trivial and easy to implement. The authors specify the hyperparameters (e.g., ridge regression lambda=10) and the calibration procedure in detail. The project page is provided, and the use of standard benchmarks and open-source models (Qwen, Voxtral, Phi-4) facilitates replication. The "screen" for identifying models where the prior fails is also described clearly.
The primary limitation is that the linear prior fails on two specific models (Qwen-Audio variants), although the authors provide a diagnostic screen to identify such cases. The method is designed for pruning before the LM runs, so it does not address KV-cache eviction during the decoding phase, which the authors note as future work. The performance gains on multiple-choice tasks are modest in absolute terms (e.g., +0.043 accuracy over DART), though statistically significant.
This work has significant practical impact for deploying audio LLMs in resource-constrained environments. By enabling longer audio inputs within fixed context windows and increasing throughput, Triage makes real-time or long-form audio understanding more feasible. The finding that attention is linearly predictable from encoder outputs is a valuable insight that could inform other efficiency techniques in multimodal models. The paper introduces a label-free, linear-predictive method for pruning audio tokens in LALMs before the language model runs, significantly improving efficiency and context window capacity. The technical contribution is strong, combining a novel insight about attention predictability with a practical, two-stage pruning algorithm that outperforms existing baselines across multiple models and tasks.
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance \tool's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of \tool for unified audio understanding across speech, music, and general audio domains.
Primary: Tsinghua University
All Institutions: Tsinghua University
The paper presents a unified audio encoder that integrates complementary encoders from Qwen2-Audio and Audio-Flamingo 3 through MoE-based fusion, achieving state-of-the-art performance on the XARES-LLM benchmark. The technical contribution is significant in demonstrating that sparse expert architectures can effectively reconcile heterogeneous audio representations from different pre-training regimes, offering a scalable and efficient path toward general-purpose audio understanding.
The paper proposes UniAE-MoE, a unified audio encoder that fuses features from two distinct Large Audio Language Model (LALM) backbones: Qwen2-Audio and Audio-Flamingo 3. The core methodological contribution is the use of a Mixture-of-Experts (MoE) architecture with SwiGLU activations to decouple and recombine these heterogeneous representations. The authors argue that while both models share a Whisper-based backbone, their distinct pre-training objectives lead to complementary feature spaces. The MoE router performs token-level conditional expert aggregation, allowing the model to specialize in different acoustic modalities (speech, music, general audio) while a shared expert captures invariant features. Additionally, the paper introduces a two-stage instruction-tuning strategy and a Task-Specific Data Scaling (TSDS) technique to address data imbalance across 20 downstream tasks. The approach is sound, leveraging existing strong encoders rather than training from scratch, which is a practical and effective strategy for current LALM development.
The experimental evaluation is robust, targeting the XARES-LLM benchmark and the Interspeech 2026 Audio Encoder Capability Challenge. The model achieves a state-of-the-art score of 0.802 on XARES-LLM, outperforming individual backbones (Qwen2-Audio: 0.737, Audio-Flamingo 3: 0.767) and simple fusion baselines like concatenation (0.778). Ablation studies confirm the importance of SwiGLU and the MoE structure. The TSDS technique shows significant gains on low-resource tasks (e.g., Free Music Archive accuracy jumping from 0.848 to 0.979). The inclusion of hidden test sets from the official challenge provides strong evidence of generalization. However, the reliance on a lightweight 135M-parameter LLM for the decoder limits the assessment of the encoder's potential with larger, more capable language models.
The authors provide a GitHub repository link, which is a positive step. The paper details the training steps (30k and 80k), hardware (3x NVIDIA L20), and specific hyperparameters (top-2 routing, hidden dimension 3,413). The data sources are clearly listed, including specific datasets for TSDS. However, the exact implementation details of the "Dasheng-base" baseline and the specific preprocessing steps for aligning the two different encoder outputs (beyond "temporal and feature-wise alignment") could be more explicit to ensure full reproducibility.
A primary limitation is the computational overhead of running two full encoder backbones (Qwen2-Audio and Audio-Flamingo 3) in parallel, which may hinder real-time or edge deployment despite the MoE efficiency. The paper acknowledges that performance on some hidden generative tasks (AISHELL-6, LibriHeavy) is lower, indicating room for improvement in robust generalization. Furthermore, the method is heavily dependent on the quality of the two specific base encoders; if one backbone is weak or biased, the fusion may not fully compensate. The use of a small LLM (SmolLM2-135M) for evaluation might underestimate the encoder's true capability in complex reasoning tasks.
This work contributes to the trend of "General Audio Intelligence" by demonstrating that combining specialized LALM encoders via MoE can yield superior unified representations. It provides a blueprint for integrating multiple audio models without retraining them from scratch, which is valuable for the community. The TSDS technique offers a practical solution for multi-task learning imbalances, applicable to other multimodal domains. The strong performance on the Interspeech 2026 challenge validates the approach in a competitive, standardized setting. The paper presents a unified audio encoder that integrates complementary encoders from Qwen2-Audio and Audio-Flamingo 3 through MoE-based fusion, achieving state-of-the-art performance on the XARES-LLM benchmark. The technical contribution is significant in demonstrating that sparse expert architectures can effectively reconcile heterogeneous audio representations from different pre-training regimes, offering a scalable and efficient path toward general-purpose audio understanding.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: MIT CSAIL
All Institutions: MIT CSAIL, MIT RLE, KAIST GSCT
The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
The paper introduces a differentiable, GPU-accelerated acoustic simulator for the vocal tract based on the linearized Euler equations. The core technical contribution is the shift from time-domain finite difference methods to a frequency-domain formulation using transmission line matrices (TLM). This allows for parallelization across frequencies, resulting in a 70x speedup and, crucially, smoother gradient landscapes that facilitate stable optimization. The authors also propose a differentiable turbulence model using softplus gating and softmax-based constriction localization to handle consonants, unifying vowel and consonant synthesis. Finally, they leverage neural fields (implicit neural representations) to parameterize the vocal tract geometry, which acts as a regularizer to escape local minima in the non-convex inverse problem.
The experiments are extensive and well-designed. The authors validate the simulator's accuracy by comparing Frequency Domain Synthesis (FDS) against Time Domain Synthesis (TDS), showing significant improvements in SI-SDR, STOI, and PESQ metrics. They demonstrate the utility of the simulator in two novel applications: (1) a self-supervised autoencoder that maps audio to vocal tract area functions across 11 languages, outperforming paired-data baselines in intelligibility and speaker identity preservation; and (2) a generative model that reconstructs MRI videos of the vocal tract from speech alone, without paired data. The emergence of the IPA vowel chart from the latent space of the MRI generator provides strong evidence for the physical grounding of the model.
The paper provides detailed derivations of the physics, including the linearized Euler equations, Webster's equation, and the circuit interpretation. Specific hyperparameters, such as window lengths, hop sizes, and Reynolds number thresholds, are provided. The use of standard architectures like Wav2Vec 2.0 and StyleGAN2 aids reproducibility. However, the specific implementation of the differentiable turbulence model and the exact neural field architectures (RFF vs MFN) would benefit from more code-level details or a public repository link beyond the project page.
The model relies on a 1D acoustic tube approximation, which may not capture all 3D effects of the vocal tract, particularly for complex articulations or nasal sounds (though the nasal tract is noted as a limitation due to MRI data availability). The inverse problem is inherently ill-posed, meaning multiple geometries can produce the same sound, leading to potential ambiguities in reconstruction. The MRI reconstruction quality is limited by the resolution and coverage of the training MRI dataset.
This work has significant implications for linguistics, speech pathology, and medical imaging. By providing a physically grounded link between speech and articulatory geometry, it enables new tools for voice coaching, language acquisition studies, and non-invasive visualization of vocal tract movements. The self-supervised nature of the autoencoder and MRI generator makes these tools scalable and accessible without the need for expensive paired data collection. The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS$^2$D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS$^2$D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS$^2$D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS$^2$D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
Primary: The University of Texas at Austin
All Institutions: The University of Texas at Austin, Shanghai Jiao Tong University
[One sentence main contribution]. The paper introduces AS$^2$D, a target-decoupled speculative decoding framework for audio language models that enables concurrent drafting and verification on mobile devices, significantly improving throughput and latency consistency compared to conventional serial speculative methods.
The paper proposes AS$^2$D (Audio Speculative Speculative Decoding), a method that decouples the drafting process from the target model's evolving verified prefix in speculative decoding. The core insight is that for source-conditioned tasks like ASR, the input audio and instruction provide sufficient context for a smaller drafter to generate candidates independently of the target's current output state. This allows drafting and verification to proceed concurrently on heterogeneous mobile processors, rather than serially. The methodology involves an audio-conditioned drafter that maintains its own generation history, publishing candidates to a buffer. The target model verifies these candidates against its own trajectory, using local token matching to align and reuse rejected candidates where possible. The approach is rigorously formulated with a problem definition that accounts for memory constraints and latency quantiles, and it includes a device-profiled configuration step to select optimal budget caps based on hardware characteristics.
The evaluation is extensive and well-suited to the mobile inference context. The authors test two target models (Qwen3-ASR-1.7B and Qwen2.5-Omni-7B) across four different Android phones with varying hardware capabilities (Snapdragon 750G to 8 Gen 3). They evaluate on seven datasets covering ASR, translation, and spoken QA, totaling 12.2 hours of audio. The results show significant throughput improvements (42-76%) over target-only decoding, with a crucial finding that AS$^2$D avoids the latency spikes common in standard speculative decoding (only 5.7% of windows are slower vs. 58-63% for baselines). The comparison against SpecASR and standard SD is fair and highlights the advantage of decoupled drafting on resource-constrained devices.
The paper provides detailed implementation notes, specifying the use of the MNN framework for Android and the specific model versions used. The description of the buffer management, alignment rules, and profiling methodology is sufficient for a skilled engineer to reproduce the system. However, the specific code for the custom MNN modifications and the profiling scripts are not explicitly linked in the provided text, which slightly limits immediate reproducibility without access to the full repository.
The primary limitation is the dependency on the specific hardware characteristics of the tested phones; the gains may vary on other architectures. The method relies on the assumption that the audio source provides enough signal for the drafter to be useful, which may not hold for all audio tasks or extremely noisy conditions. Additionally, the approach is tailored to autoregressive generation; it does not apply to non-autoregressive models. The paper focuses on throughput and latency but does not extensively discuss the energy consumption trade-offs of running two models concurrently, which is critical for mobile devices.
This work has significant implications for the deployment of large audio language models on edge devices. By enabling concurrent drafting and verification, it makes high-quality, real-time audio understanding feasible on smartphones without cloud dependency. This could enable new applications in privacy-preserving voice assistants, wearable memory aids, and offline translation tools. The technique of decoupling drafting from the target prefix could potentially be adapted to other source-conditioned generation tasks, such as video captioning or code completion with context. [One sentence main contribution]. The paper introduces AS$^2$D, a target-decoupled speculative decoding framework for audio language models that enables concurrent drafting and verification on mobile devices, significantly improving throughput and latency consistency compared to conventional serial speculative methods.
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Primary: Georgia Institute of Technology
All Institutions: Georgia Institute of Technology, Dolby Laboratories
The paper introduces a training-free agentic framework that integrates spatially aware sound generation into 3D world models by explicitly linking semantic labels, reconstructed geometry, and geometric acoustic propagation. This approach effectively solves the problem of spatial inconsistency in generated audio by grounding sound sources in a persistent 3D world state, resulting in significant improvements in spatial consistency and audio-visual alignment compared to existing video- and text-conditioned baselines.
The paper proposes a training-free, agentic pipeline that bridges visual world generation and spatial audio synthesis. The core methodological contribution is the formulation of an "audible world state" that explicitly separates semantic sound labels, dry audio assets, 3D source placements, and acoustic parameters. By leveraging VLMs to parse panoramic 3D proxies into semantic layers (foreground objects vs. ambient background), the system grounds generated audio to reconstructed geometry. It then utilizes a geometric acoustic simulator (GSound) to render listener-dependent binaural audio. This approach is technically sound and effectively addresses the lack of spatial persistence in current video-to-audio models. The use of an agentic framework for parameterization and source grounding is a clever engineering solution that avoids the need for end-to-end training of complex multimodal models.
The evaluation is comprehensive, covering both semantic alignment and spatial consistency. The authors utilize 80 generated scenes and compare against strong baselines including Stable Audio 1.0, MMAudio, SEE-2-SOUND, and OmniAudio. The use of Direction of Arrival (DoA) MAE and F1 scores for spatial consistency is a rigorous metric that directly tests the core claim of the paper. The inclusion of human evaluations for preference and VLM-based assessments adds depth to the subjective quality assessment. The results show substantial gains in spatial consistency (6.91ยฐ DoA MAE for isolated sources) while maintaining competitive semantic scores, validating the effectiveness of the proposed framework.
The paper provides a clear description of the pipeline stages and the specific tools used (HunyuanWorld, Stable Audio 1.0, GSound). However, as a training-free system relying on specific API calls and agentic prompts, full reproducibility depends on the stability of these third-party models and the specific prompt templates (referenced in appendices). The lack of a public code repository or demo link in the provided text slightly hinders immediate reproducibility, though the modular nature of the system makes it feasible to replicate with the cited components.
The system relies on the quality of the initial panoramic generation and depth estimation; errors in the 3D proxy will propagate to source placement. The acoustic parameterization is heuristic and "conservative," meaning it may not capture complex material-specific acoustic properties accurately. Additionally, the pipeline is computationally intensive due to the multi-stage agentic process and geometric simulation, which may limit real-time applications. The evaluation is limited to 80 scenes, which, while sufficient for a proof of concept, is relatively small for a general claim about world models.
This work has significant implications for the development of immersive virtual environments, VR/AR applications, and interactive storytelling. By making audio a persistent, geometry-aware property of the world state rather than a post-hoc soundtrack, it paves the way for more consistent and engaging user experiences. The framework's modularity allows for easy integration of future improvements in audio generation or 3D reconstruction, making it a robust foundation for future research in multimodal world modeling. The paper introduces a training-free agentic framework that integrates spatially aware sound generation into 3D world models by explicitly linking semantic labels, reconstructed geometry, and geometric acoustic propagation. This approach effectively solves the problem of spatial inconsistency in generated audio by grounding sound sources in a persistent 3D world state, resulting in significant improvements in spatial consistency and audio-visual alignment compared to existing video- and text-conditioned baselines.
Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose SENSE(Semantic-EEG Neural Speech SynthEsis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.
Primary: Chung-Ang University
All Institutions: Chung-Ang University, University of Birmingham
SENSE reformulates EEG-to-speech synthesis by integrating spatial graph encoding with paradigm-aware semantic conditioning, demonstrating that leveraging the N400 semantic structure significantly improves both acoustic fidelity and semantic consistency, particularly in low-data cross-subject scenarios.
The paper proposes SENSE, a framework for EEG-to-speech synthesis that addresses two specific gaps in prior work: the lack of spatial modeling of EEG electrodes and the neglect of semantic structure in the N400 paradigm. The methodology is technically sound and well-motivated. The use of a graph-based encoder with a fixed anatomy-derived adjacency matrix is a logical choice to capture spatial correlations without overfitting on the limited dataset size (7,200 pairs). The integration of EEG Semantic Conditioning (ESC) is the core novelty; by aligning EEG representations with a frozen CLIP space using only congruent trials, the authors effectively leverage the known neurophysiological properties of the N400 component. The training strategy, which uses a bridge mechanism to expose the decoder to EEG-derived conditioning vectors during training, is a clever solution to the distribution shift problem at inference time. The use of S4 blocks for temporal encoding is consistent with recent state-of-the-art baselines like FESDE, ensuring a fair comparison.
The experiments are conducted on the standard N400 dataset. The paper claims consistent outperformance on both acoustic (MCD, Mel-Correlation) and semantic (WER) metrics. A particularly strong result is the generalization capability: the model trained on only two subjects surpasses the strongest baseline trained on all 18 subjects in Word Error Rate. This suggests that the semantic conditioning effectively captures invariant features across subjects, which is a significant practical advantage for BCI applications where per-subject calibration is burdensome. The ablation studies likely confirm the contribution of the graph encoder and the ESC module, though the truncated text prevents a full review of the ablation tables. The channel attribution analysis linking model reliance to auditory and sensorimotor regions adds biological plausibility to the technical results.
The authors provide a link to a project page with code and audio samples. The methodology section provides sufficient detail on the graph construction (BioSemi montage, Gaussian kernel), the encoder architecture (S4 blocks, Conformer for phonemes), and the loss functions. The specific hyperparameters for the graph convolution and the initialization of the channel gates are mentioned. The use of standard components (VITS, CLIP ViT-B/32, HiFi-GAN) further aids reproducibility.
The primary limitation is the reliance on the N400 paradigm, which is a passive listening task. It is unclear how well this approach generalizes to active speech production or different speech contexts. The use of a fixed adjacency matrix, while justified by data scarcity, may limit the model's ability to adapt to individual variations in brain topology. Additionally, the semantic alignment is restricted to congruent trials, which means the model may not effectively handle or distinguish incongruent semantic processing in a generative context, although this is arguably outside the scope of standard speech synthesis.
This work has significant implications for non-invasive BCIs aimed at restoring communication for individuals with conditions like ALS or locked-in syndrome. By demonstrating that semantic information can be effectively extracted from EEG and used to condition speech synthesis, the paper moves the field beyond mere acoustic reconstruction towards meaningful communication. The strong cross-subject generalization results are particularly impactful, as they suggest that large-scale per-subject training may not be necessary, making the technology more accessible and practical for real-world deployment. SENSE reformulates EEG-to-speech synthesis by integrating spatial graph encoding with paradigm-aware semantic conditioning, demonstrating that leveraging the N400 semantic structure significantly improves both acoustic fidelity and semantic consistency, particularly in low-data cross-subject scenarios.
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
Primary: Meta AI (Reality Labs / FAIR)
All Institutions: Reality Labs at Meta, FAIR at Meta, National Taiwan University
The paper introduces EmoRES, a training-free vector steering method that decomposes emotion vectors into shared and residual components to enhance emotional control in TTS. By independently scaling the shared (neutral-to-centroid) and residual (centroid-to-emotion) components, EmoRES improves emotional accuracy and naturalness over conventional steering methods without retraining the backbone, demonstrating significant gains in objective metrics and human evaluation across multiple state-of-the-art TTS systems.
The paper proposes EmoRES, a training-free vector steering method for emotional TTS. The core contribution is the decomposition of the standard mean-difference emotion vector ($v_e = h_e - h_{neu}$) into a shared component (neutral to centroid) and a residual component (centroid to specific emotion). By independently scaling these components ($\alpha_c$ and $\alpha_r$), the method allows for enhanced categorical specificity without amplifying the generic "emotional" shift. The methodology is mathematically sound, leveraging the linearity of the steering vectors in the activation space. The approach is elegant in its simplicity, requiring no retraining or new subspace learning, and effectively addresses the limitation of conventional steering where a single global strength couples generic emotional intensity with specific emotional identity.
The experiments are rigorous, utilizing two distinct state-of-the-art frozen backbones (IndexTTS-2 and CosyVoice2) to demonstrate generality. The evaluation protocol is strong, employing a held-out test set (IEMOCAP) distinct from the development set (CREMA-D) to ensure out-of-distribution generalization. The metrics cover objective emotion recognition (TEP, E-SIM, Rank Correlation, Hit Rate), speaker similarity, and intelligibility (WER). Crucially, the paper includes human evaluation for both emotion identification accuracy and naturalness preference, providing a holistic view of performance. The results show consistent improvements over the CoCoEmo baseline, particularly in rank correlation and hit rate, validating the hypothesis that residual enhancement improves emotional control.
The paper provides sufficient detail for reproduction, including the specific layers used for steering, the datasets for vector extraction (ESD, CREMA-D, RAVDESS), and the normalization procedure (rescaling to mean Euclidean norm). The decomposition formula is clearly defined. However, the code is not explicitly linked in the provided text, and specific hyperparameters for the "signal-quality gate" and "speech-emotion-recognition gate" during vector extraction are deferred to the appendix, which may limit immediate reproducibility without access to those details.
The method relies on the assumption that the centroid of emotional activations is a meaningful functional boundary between "neutral" and "emotional" states, which may not hold for all emotion categories or speakers. The evaluation is limited to four discrete emotions (angry, happy, sad, surprise); performance on more nuanced or continuous emotional dimensions is unknown. Additionally, while the method is training-free, it requires a curated library of paired neutral-emotional utterances for vector extraction, which may be a bottleneck for rare emotions.
This work contributes to the growing field of interpretability and control in generative speech models. By demonstrating that internal representations of emotion have a decomposable structure, it offers insights into how TTS models encode affective information. The training-free nature of the approach makes it highly practical for deploying emotional control in existing commercial or open-source TTS systems without the cost of retraining. It also sets a baseline for future work on fine-grained emotional control and mixed-emotion synthesis. The paper introduces EmoRES, a training-free vector steering method that decomposes emotion vectors into shared and residual components to enhance emotional control in TTS. By independently scaling the shared (neutral-to-centroid) and residual (centroid-to-emotion) components, EmoRES improves emotional accuracy and naturalness over conventional steering methods without retraining the backbone, demonstrating significant gains in objective metrics and human evaluation across multiple state-of-the-art TTS systems.
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $ฯ$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $ฯ$-Voice with an acceptable increase in the delegation rate.
Primary: Tsinghua University
All Institutions: Tsinghua University, ByteDance
SALMONN-duo introduces an adaptive dual-system voice agent that balances real-time responsiveness with deep reasoning via knowledge-boundary-aware training and cost-aware reinforcement learning. The paper presents a rigorous methodology for training a full-duplex frontend to delegate tasks to an asynchronous backend only when necessary, demonstrating superior performance-cost trade-offs and improved safety in complex, multi-turn, environment-grounded tasks compared to existing full-duplex voice agents.
The paper proposes SALMONN-duo, a dual-system architecture for full-duplex voice agents. The core methodological contribution is the separation of a lightweight, always-on full-duplex frontend (System 1) from a powerful, asynchronous backend agent (System 2). The novelty lies in the "adaptive delegation" mechanism, where System 1 learns to decide when to answer directly versus when to delegate to System 2. This is achieved through two key training strategies: (1) Knowledge-Boundary-Aware Supervised Fine-Tuning (SFT), which uses the model's own rollout correctness to label delegation targets, and (2) Cost-Aware Group Relative Policy Optimization (GRPO), which introduces a reward function that balances task success, safety (hallucination/policy violation), and the cost of unnecessary delegation. The design is practical, addressing the latency-cost trade-off inherent in real-time voice interactions. The use of a delegation token within the response stream rather than a separate head is a specific architectural choice validated in the paper.
The experimental evaluation is comprehensive, covering single-turn QA, multi-turn daily conversations, and complex environment-grounded tasks (using a customized $\tau$-Voice benchmark). The paper compares against recent full-duplex and half-duplex baselines (Moshi, VoiceChat, etc.). Results show that SALMONN-duo achieves a better performance-cost trade-off, with lower invocation rates on easier tasks and higher accuracy on complex reasoning tasks compared to baselines that either always invoke or rarely invoke tools. The inclusion of safety rewards in the RL phase demonstrates a reduction in hallucinations and policy violations. The evaluation of interruption handling (barge-in) is a strong point, showing the model can maintain conversational context while waiting for backend responses.
The paper provides detailed descriptions of the data generation pipeline, including the roles of user, assistant, and supervisor simulators. It specifies model architectures (Llama-3.1-8B, CosyVoice2), training hyperparameters (learning rates, batch sizes, GPU counts), and reward weights. However, the reliance on proprietary models (GPT-4o, GPT-5.2, OpenAI TTS) for simulation and evaluation limits full open-source reproducibility. The code and specific weights are not explicitly linked in the provided text, though the paper claims to be open-source work.
The system relies on a strong proprietary backend (GPT-5.2) for the "System 2" component, which may not be accessible to all researchers or deployable in privacy-sensitive on-device scenarios. The evaluation assumes ground-truth dialogue history for the backend in some settings, which may overestimate performance in noisy real-world ASR conditions (though ASR sensitivity is analyzed). The "cost" is defined primarily as the number of backend invocations, not necessarily computational cost or latency in a strict sense, though latency is simulated.
This work is significant for the development of practical, low-latency voice assistants that can handle complex tasks without sacrificing conversational fluidity. The dual-process approach offers a scalable path for integrating heavy reasoning capabilities into real-time speech interfaces. It addresses a critical gap in current voice agents: the inability to efficiently manage the trade-off between responsiveness and deep reasoning/tool use. SALMONN-duo introduces an adaptive dual-system voice agent that balances real-time responsiveness with deep reasoning via knowledge-boundary-aware training and cost-aware reinforcement learning. The paper presents a rigorous methodology for training a full-duplex frontend to delegate tasks to an asynchronous backend only when necessary, demonstrating superior performance-cost trade-offs and improved safety in complex, multi-turn, environment-grounded tasks compared to existing full-duplex voice agents.
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.
Primary: The Hong Kong University of Science and Technology (Guangzhou)
All Institutions: The Hong Kong University of Science and Technology (Guangzhou)
The paper presents a novel training-free framework for speech emotion editing by leveraging dynamic velocity transport in flow-matching TTS models, supported by a systematic analysis of editability and a new benchmark. The technical contribution is strong, offering a robust alternative to training-based and static steering methods, with comprehensive experiments validating its effectiveness across multiple state-of-the-art backbones.
The paper proposes SEmoEdit, a training-free framework for speech emotion editing based on flow-matching TTS models. The core contribution is formulating emotion editing as dynamic velocity transport. Instead of using static activation steering vectors (which are often unstable and require paired data to derive), the method dynamically computes the velocity difference between source and target conditions at each step of the ODE integration. This approach is theoretically grounded in the properties of Conditional Flow Matching (CFM). The method also introduces "emotion bridging" to handle interpolation artifacts, where intermediate states are not well-formed speech samples. The inclusion of a systematic probing study (Sec. 2) to diagnose *when* and *how* editing works along the trajectory is a strong methodological addition, providing empirical justification for the design choices (e.g., skipping early steps to preserve timing).
The authors introduce SEmoEditBench, a 600-case benchmark, which is a valuable contribution to the field as standardized benchmarks for speech editing are scarce. Experiments are conducted on three SOTA backbones (F5-TTS, CosyVoice 2, IndexTTS 2), demonstrating broad applicability. The evaluation metrics are comprehensive, covering effectiveness (TEP, SES), preservation (WER, S-SIM), quality (UTMOS), and subjective evaluation (MOS). The results show that SEmoEdit outperforms training-based and activation-steering baselines. The ablation studies on trajectory timing and attribute factorization are insightful and support the paper's claims.
The paper provides a GitHub link with code, benchmark, and audio samples. The methodology is described in sufficient detail (equations for velocity transport, noise coupling, and bridging) to allow reproduction. The use of standard open-source TTS models (F5-TTS, CosyVoice 2) further enhances reproducibility.
The method relies on the quality of the underlying TTS model's velocity field; if the base model is poor at emotion generation, the editing will suffer. The "emotion bridging" step for interpolation adds computational overhead and complexity. The benchmark, while useful, is relatively small (600 cases) compared to large-scale speech datasets. The paper focuses on English and Chinese (implied by ESD/IEMOCAP/RAVDESS/CREMA-D usage), and generalization to other languages is not explicitly tested.
This work has significant implications for the controllability of generative speech models. By demonstrating that pretrained models contain latent editing capabilities that can be unlocked via inference-time manipulation, it opens the door to more flexible and efficient speech editing pipelines without the need for fine-tuning. This is particularly relevant for applications in dubbing, accessibility, and personalized voice assistants where emotional nuance is important. The insights into trajectory-dependent editability can guide the design of future controllable TTS systems. The paper presents a novel training-free framework for speech emotion editing by leveraging dynamic velocity transport in flow-matching TTS models, supported by a systematic analysis of editability and a new benchmark. The technical contribution is strong, offering a robust alternative to training-based and static steering methods, with comprehensive experiments validating its effectiveness across multiple state-of-the-art backbones.
Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like hallucination. To address this modality-gap issue, we introduce JELAR, a Joint-Embedding Predictive Architecture (JEPA)-based latent reasoning framework that conditions latent reasoning supervision on acoustic representations learned from raw waveforms. During training, a frozen WavJEPA model provides representations learned directly from raw waveforms. A non-causal expert first constructs answer-aware queries, which cross-attend to WavJEPA embeddings to produce latent reasoning targets. The LALM is trained to predict these targets before generating its response. Experimental results show that JELAR improves the Audio-Reasoner baseline by 2.70 and 9.10 absolute percentage points on MMAU-mini and MMAR, respectively, demonstrating the effectiveness of JEPA-conditioned latent reasoning as an alternative to explicit textual CoT supervision.
Primary: Nanyang Technological University
All Institutions: Nanyang Technological University, AI Singapore, Peking University
[One sentence main contribution]. The paper introduces JELAR, a JEPA-conditioned latent reasoning framework that improves audio reasoning accuracy by grounding latent supervision in raw waveform representations, effectively addressing the modality gap inherent in textual Chain-of-Thought approaches.
The paper proposes JELAR, a framework that replaces explicit textual Chain-of-Thought (CoT) with latent reasoning conditioned on acoustic representations. The core methodological contribution is the use of a frozen WavJEPA encoder to generate acoustic embeddings, which are then processed by a non-causal expert (BERT-based) to create answer-aware latent targets. These targets serve as supervision for the LALM during training. The approach is logically sound, addressing the "modality gap" where text-based CoT may hallucinate or miss acoustic details. However, the reliance on a separate, non-causal expert for target generation introduces complexity and a potential train-test mismatch (teacher forcing vs. autoregressive inference), which is a known challenge in latent reasoning. The integration of JEPA is a creative application of predictive coding principles to audio reasoning.
The experiments are conducted on MMAU-mini and MMAR benchmarks, comparing JELAR against the Audio-Reasoner baseline. The reported improvements are significant, particularly on MMAR (+9.10 points), suggesting the method is effective for multi-step reasoning. Ablation studies confirm the contribution of acoustic conditioning and latent reasoning length. However, the evaluation is limited to these two benchmarks, and comparisons with other state-of-the-art audio LLMs (like GPT-4o or Gemini) are not directly made in the main results table for the proposed method, only for context. The lack of qualitative analysis of the latent reasoning process is a minor weakness.
The paper provides sufficient detail on the architecture and training procedure, including the use of frozen WavJEPA and BERT-base. Hyperparameters like K=20 are specified. However, specific details on the training data composition, optimization hyperparameters (learning rates, batch sizes), and the exact implementation of the cross-attention module are not fully detailed in the provided text. The code repository is not linked, which hinders immediate reproducibility.
The method requires a separate non-causal expert during training, which increases computational overhead and complexity. The latent reasoning space is not interpretable, making it difficult to debug errors. The performance gains, while positive, are based on a specific baseline (Audio-Reasoner) and may not generalize to other LALM architectures without further adaptation. The paper does not discuss the computational cost of inference compared to standard CoT or direct answering.
This work contributes to the growing field of latent reasoning in multimodal models, offering an alternative to text-centric CoT. It highlights the utility of JEPA-style representations for audio tasks, potentially inspiring similar approaches in other modalities. The focus on reducing hallucination by grounding reasoning in acoustic evidence is a valuable direction for improving the reliability of audio LLMs. [One sentence main contribution]. The paper introduces JELAR, a JEPA-conditioned latent reasoning framework that improves audio reasoning accuracy by grounding latent supervision in raw waveform representations, effectively addressing the modality gap inherent in textual Chain-of-Thought approaches.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: HKGAI
All Institutions: HKGAI
YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The paper proposes a unified architecture (YuE2) that bridges symbolic and audio music generation using an AR-NAR Mixture-of-Transformers (MoT). The core methodological contribution is the "symbolic planning" stage, where the model first generates a readable score (melody, harmony, form) before expanding into semantic tokens and finally acoustic latents via flow matching. This hierarchical approach allows for explicit control over composition. Additionally, the paper introduces two auxiliary models: MERT2 for semantic representation learning and SheetSage2 for symbolic supervision (lead-sheet transcription), which are used to train the main model on unaligned audio data. The integration of these components into a single checkpoint that supports editing and cover generation is a significant architectural advancement.
The evaluation is extensive, utilizing both objective benchmarks (WildSongBench, SongBench, MARBLE) and subjective expert listening tests. The results show YuE2 outperforming public baselines and being competitive with proprietary systems like Suno v4.5 and v5. The ablation study comparing generation with and without symbolic planning demonstrates a clear preference for the planned approach (49.3% vs 34.6%). The introduction of MERT2 and SheetSage2 is supported by state-of-the-art results on their respective benchmarks (MARBLE and transcription metrics). The inclusion of zero-shot cover generation and agentic editing case studies further validates the utility of the unified framework.
As a technical report from an industry lab (HKGAI), the paper likely lacks full open-source code and model weights at the time of release, which limits immediate reproducibility. However, the detailed description of the architecture (AR-NAR MoT, flow matching) and the specific benchmarks used provides a clear roadmap for replication. The reliance on proprietary datasets for training the auxiliary models may pose a barrier for independent researchers.
The paper relies heavily on expert listening tests, which can be subjective and difficult to scale. The comparison with proprietary systems (Suno) is limited to specific versions and may not reflect the full capability of those closed-source models. The "best-of-8" selection strategy, while effective for benchmarking, introduces a computational overhead that may not be practical for real-time applications. Additionally, the paper does not extensively discuss the computational cost of the multi-stage generation process compared to direct audio generation.
This work has significant implications for the music production industry by providing a tool that allows for both high-quality audio generation and explicit compositional control. The ability to edit scores and generate covers zero-shot opens up new possibilities for music creation, education, and collaboration. The unification of symbolic and audio domains could lead to more interpretable and controllable generative music systems, potentially influencing how AI is integrated into professional music workflows. YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Primary: The Chinese University of Hong Kong
All Institutions: The Chinese University of Hong Kong, Microsoft Corporation
The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
The paper proposes DuraS2ST, a framework that reformulates duration-aligned speech-to-speech translation (S2ST) as a reasoning problem rather than an acoustic post-processing task. The core methodological contribution is the integration of Chain-of-Thought (CoT) reasoning into the speech generation pipeline. The model first generates an explicit textual rationale planning the target wording and phonetic length, followed by the synthesis of interleaved text-acoustic tokens. This is supported by the construction of DuraSet-440K, a large-scale corpus specifically designed for this purpose. The training paradigm is two-phase: Supervised Fine-Tuning (SFT) on the new corpus, followed by Group Relative Policy Optimization (GRPO). The RL phase introduces two specific technical innovations: the Duration Margin Reward (DMR), which acts as a soft constraint to balance translation quality with duration consistency, and Modality-Aware Reward Attribution (MARA), which prevents cross-modality reward contamination by assigning duration rewards only to acoustic tokens and quality rewards to both text and acoustic spans. This approach is technically sound and addresses a specific gap in current S2ST systems where duration control is often handled via black-box speed tokens or post-hoc time-stretching.
The experiments are conducted on the CVSS-T benchmark, which is disjoint from the training data, providing a valid out-of-domain evaluation. The paper compares DuraS2ST against strong baselines, including commercial models (GPT-4o, Qwen2.5-Omni, Kimi-Audio) and open-source SLMs (Step-Audio-2-mini). The results demonstrate that DuraS2ST achieves a superior balance between translation quality (BLEU, COMET) and duration consistency (SLC-0.2, SLC-0.4, MADE, MRDE). Notably, the method shows substantial gains in a zero-shot RL setting, highlighting the effectiveness of the reward design. The inclusion of both open-source and commercial baselines strengthens the experimental validation.
The paper provides a project page with a GitHub repository link. It details the training setup, including hyperparameters for SFT and GRPO, and describes the construction pipeline for DuraSet-440K in detail. The use of standard metrics (BLEU, COMET, SECS, SLC) and a public benchmark (CVSS-T) enhances reproducibility. However, the specific implementation details of the "duration-controllable TTS model" used for data synthesis are not fully specified in the main text, though likely detailed in the appendix or code.
The method relies on a specific base model (Step-Audio-2-mini-Think), and its generalizability to other SLM architectures is not extensively tested. The construction of DuraSet-440K is resource-intensive, requiring LLM-based translation candidate generation and duration-controllable TTS synthesis. The evaluation is limited to English-Chinese pairs; performance on other language pairs with different phonetic structures remains unknown. Additionally, the "zero-shot RL" claim needs careful interpretation as it likely refers to the RL phase without additional SFT on the target domain, but the model is still SFT'd on DuraSet-440K.
This work has significant implications for video dubbing, simultaneous interpretation, and real-time communication applications where audio-visual synchronization is critical. By treating duration alignment as a reasoning task, it offers a more interpretable and controllable approach compared to traditional acoustic constraints. The release of DuraSet-440K and the code will likely accelerate research in duration-controlled speech generation and multimodal reasoning. The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
Primary: Alibaba Group
All Institutions: Shanghai Jiao Tong University, Alibaba Group, Tsinghua University, University of Cambridge, Nankai University, Chinese University of Hong Kong, SII
Pruned CTC enables memory-efficient training of large-vocabulary ASR models by exploiting the sparsity of CTC alignments, and LLM-CTC successfully adapts pretrained LLMs for fast, non-autoregressive speech recognition with strong performance in both offline and streaming settings.
The paper introduces "Pruned CTC," a memory-efficient implementation of Connectionist Temporal Classification (CTC) that exploits the sparsity of valid alignments. The core insight is that while the softmax normalization requires the full vocabulary, the dynamic programming alignment only involves the target tokens and the blank symbol. By restricting the alignment computation to this small subset while retaining full-vocabulary normalization via a "complement" class, the authors prove exact equivalence in loss and gradients. This is a mathematically sound and clever optimization. The method is further extended to "LLM-CTC," which adapts pretrained Large Language Models (LLMs) for non-autoregressive ASR. This is a significant architectural contribution, as it allows leveraging the linguistic knowledge of LLMs without the latency of autoregressive decoding, and it handles streaming via bounded-history attention masks without requiring chunk-level alignments.
The experiments are robust and cover a wide range of model sizes (0.6B to 32B Qwen3 models) and datasets (GigaSpeech, etc.). The results demonstrate a 5.1x reduction in memory with minimal time overhead. The comparison against LLM-CE (Cross-Entropy) shows that LLM-CTC achieves comparable accuracy (within 7% relative WER) while being 7-10x faster in recognition. The streaming results are particularly strong, showing that the method avoids the complex alignment requirements of previous streaming LLM-ASR methods. The inclusion of both offline and streaming evaluations provides a comprehensive view of the method's utility.
The authors provide a GitHub repository link and state that code and pretrained models will be open-sourced. The paper includes detailed algorithmic descriptions (Algorithm 1) and mathematical proofs, which aids reproducibility. The specific hyperparameters for the alignment beam and chunking are provided in the appendices (referenced).
The method relies on the assumption that the union of target tokens in a batch is small relative to the full vocabulary, which holds for typical ASR tasks but might not for extremely diverse or noisy data. The "finite-beam" approximation introduces a small error, though the paper claims it is negligible ($10^{-11}$). The streaming extension requires careful management of KV caches and attention masks, which can be complex to implement in production systems.
This work has high potential impact on the deployment of LLM-based ASR systems. By enabling non-autoregressive decoding with LLMs, it significantly reduces inference latency and computational cost, making real-time, high-quality speech recognition more accessible. The memory efficiency also allows for training larger models or larger batches on standard hardware. This could accelerate the adoption of LLMs in speech processing pipelines. Pruned CTC enables memory-efficient training of large-vocabulary ASR models by exploiting the sparsity of CTC alignments, and LLM-CTC successfully adapts pretrained LLMs for fast, non-autoregressive speech recognition with strong performance in both offline and streaming settings.
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing, REtokenization-Durable Watermarking IN Generation. It builds a graph from the token substitutions observed under retokenization, whose Laplacian yields a basis that assigns similar values to tokens likely to substitute for one another. Over this basis, embedding and detection functions are jointly optimized to preserve watermark signal through retokenization while limiting embedding distortion and detector variability on unwatermarked speech. On the Moshi full-duplex system, after eight consecutive passes of Mimi resynthesis, Redwing achieves 80.7% TPR at a calibrated 1% FPR, compared with 8.3% for KGW and at most 7.3% for WMAR. It also has the highest TPR after eight passes through three other neural codecs (77.5-93.0%), and the gains generalize to TTS models at a speech-quality cost close to that of KGW. These results show that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.
Primary: University of Zurich and ETH Zurich
All Institutions: Institute of Neuroinformatics, NAVER Cloud, University of Zurich, ETH Zurich
[One sentence main contribution]. The paper introduces Redwing, a spectral graph-based token watermarking method that exploits the structure of codec retokenization to achieve high robustness against repeated resynthesis attacks, significantly outperforming existing token-level and post-hoc watermarking techniques in speech generation systems.
The paper proposes "Redwing," a token-level watermarking scheme for speech generation that leverages the structural properties of neural codec retokenization. Unlike standard token watermarks (like KGW) that rely on specific token identities, which are often altered during the decode-encode cycle, Redwing constructs a graph where nodes are tokens and edges represent substitution frequencies observed during retokenization. By using the eigenvectors of the graph Laplacian (specifically those with small eigenvalues) as a basis, the method ensures that tokens likely to substitute for one another have similar watermark values. The embedding and detection functions are jointly optimized over this basis to maximize robustness against resynthesis while minimizing distortion to the generated speech. This is a clever application of spectral graph theory to a practical audio security problem, moving beyond treating retokenization as noise to exploiting its deterministic structure.
The experiments are rigorous, testing the watermark on the Moshi full-duplex dialogue system and two TTS models (CosyVoice3, MOSS-TTS). The primary attack simulated is repeated resynthesis through the Mimi codec (8 passes). Redwing achieves an 80.7% True Positive Rate (TPR) at 1% False Positive Rate (FPR), significantly outperforming baselines like KGW (8.3%) and WMAR (7.3%). The paper also demonstrates generalizability to other neural codecs (77.5-93.0% TPR) and maintains speech quality comparable to the KGW baseline. The inclusion of multiple models and codecs strengthens the claim of general applicability.
The method is described as "training-free" in the sense that it does not require retraining the speech model, but it does require building the substitution graph from a corpus and solving for embedding/detection functions. The paper states it applies to released models unchanged, which is a strong reproducibility feature. However, specific details on the optimization process for the embedding/detection functions and the exact corpus used for graph construction are not fully detailed in the truncated text, though the methodology is clearly defined.
The method relies on the stability of the retokenization substitution patterns. If a codec is updated or if the attack involves a different type of processing (e.g., heavy noise addition, pitch shifting) that disrupts the token substitution graph, the watermark's robustness may degrade. Additionally, the construction of the graph requires access to the codec's encoder/decoder to perform the substitution counting, which might not be feasible for all black-box systems.
This work has significant implications for the provenance of AI-generated speech. As speech models become more capable and widespread, robust watermarking is essential for distinguishing synthetic from human speech. By addressing the specific weakness of token-level watermarks in audio (retokenization), this paper provides a viable path for deploying watermarks in real-world speech generation systems without modifying the underlying model architecture. [One sentence main contribution]. The paper introduces Redwing, a spectral graph-based token watermarking method that exploits the structure of codec retokenization to achieve high robustness against repeated resynthesis attacks, significantly outperforming existing token-level and post-hoc watermarking techniques in speech generation systems.
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
Primary: Sapienza University of Rome
All Institutions: Sapienza University of Rome, Moises Systems, Inc., Paradigma
SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
The paper proposes SAGE, a 105M parameter Variational Autoencoder (VAE) for audio that operates on the complex Short-Time Fourier Transform (STFT) rather than raw waveforms or magnitude spectrograms. The core architectural innovation is the adaptation of the SwinV2 vision transformer backbone to audio, utilizing rectangular patches and attention windows to handle the time-frequency structure of spectrograms. A key methodological contribution is the "semantic distillation" strategy, where the latent space is aligned with embeddings from a frozen CLAP (Contrastive Language-Audio Pretraining) model. This alignment is introduced via a delayed schedule (detached warm-up) to ensure reconstruction quality is established before semantic constraints are applied, preventing the semantic loss from degrading fidelity. The training process is divided into two phases: pretraining with a full adversarial objective (using a WavTokenizer discriminator) and a fine-tuning phase where the encoder is frozen and the decoder is refined. The use of a sum-and-difference loss for stereo imaging is also a notable technical detail, addressing the common issue of stereo collapse in autoencoders.
The experimental evaluation is rigorous and comprehensive. The authors compare SAGE against five strong baselines (Stable Audio Open, SAME-L, SAME-S, CoDiCodec, Music2Latent) across five different held-out datasets, including in-domain (FMA) and out-of-domain (MoisesDB, MusicCaps, Song Describer) sets. Metrics cover perceptual (CLAP similarity), distributional (Frรฉchet Audio Distance in MERT, PANN, and CLAP spaces), and sample-exact (SI-SDR, STFT distance) categories. Crucially, the paper includes a MUSHRA listening test with 21 raters, showing SAGE is statistically indistinguishable from the much larger SAME-L model. The semantic evaluation is extensive, using 19 probing tasks (genre, artist, instrument, etc.) in the MAEB format, where SAGE achieves state-of-the-art results on all tasks. The ablation studies on the semantic distillation schedule and adversarial components provide strong evidence for the design choices.
The paper offers high reproducibility. The authors provide a link to the GitHub repository containing code, weights, and the evaluation harness. The training data consists of publicly available corpora (FMA, MTG-Jamendo, M4Singer), and the paper details the specific splits and preprocessing steps. The hyperparameters for both training phases are listed in the appendix. The use of standard libraries and the release of the evaluation harness allows other researchers to easily verify the results and compare their own models.
The primary limitation is the reliance on the CLAP model for semantic alignment; if the CLAP embeddings are biased or limited in their musical understanding, SAGE's latent space will inherit these limitations. The model is trained on music, so its performance on non-music audio (speech, sound effects) is not evaluated and likely poor. The inference cost, while competitive, is still higher than simple convolutional codecs, which may be a barrier for real-time mobile applications. The paper focuses on music, so generalization to other audio domains is an open question.
SAGE provides a lightweight, high-fidelity, and semantically rich audio encoder that can serve as a drop-in replacement for existing autoencoders in latent diffusion pipelines for music generation. Its semantic structure could enable better control and conditioning in generative models. The efficient inference cost makes it suitable for interactive applications. The comprehensive benchmarking harness contributes to the standardization of audio autoencoder evaluation. SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Primary: Ca' Foscari University of Venice
All Institutions: Ca' Foscari University of Venice, Kandinsky Lab, Sleeping AI, Singapore University of Technology and Design
The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
The paper proposes DEFINE, a framework for disentangling speaker identity and accent in zero-shot TTS. The core technical contribution is the "prototype anchoring" mechanism, which uses a learned table of accent prototypes to supervise an exemplar encoder (based on frozen XLS-R features). This addresses the weak supervision signal for accent representations in flow-matching objectives. The method injects the accent embedding into the F5-TTS backbone via additive shifts to timestep and text embeddings, scaled by RMS to ensure proper magnitude. A key feature is the inference-time guidance weight $w$ that allows continuous control over accent strength without retraining. The use of LoRA for parameter-efficient adaptation is standard but effective. The methodology is sound, though the reliance on a specific backbone (F5-TTS) and the specific injection point (adaLN modulation via embedding shifts) limits generalizability to other architectures.
The experiments are comprehensive, evaluating seen, held-out, and out-of-domain accents. The use of a 15-way logistic regression probe on XLS-R features for accent accuracy is a reasonable objective metric, though it is a proxy for perceptual accent. The comparison against a Seed-VC cascade is strong, showing that DEFINE matches accent transfer performance while improving speaker similarity. The listening test with 11 listeners is a positive addition, confirming that the accent changes are perceptually real and that the guidance weight controls the perceived accent strength. However, the sample size for the listening test (11 listeners, 10 trials) is relatively small, and the statistical significance, while present, is based on a limited number of votes. The WER and UTMOS metrics are standard and reported, showing that intelligibility and quality are maintained.
The paper provides a GitHub link, which is a strong indicator of reproducibility. The training data sources (Common Voice, in-house studio recordings) are described, but the in-house data is not publicly available, which limits full reproducibility. The hyperparameters for LoRA, the prototype table size, and the guidance weight are specified. The use of frozen XLS-R and specific layer selection (layer 15) is detailed. Overall, reproducibility is good, assuming access to the code and the ability to recreate the in-house speaker prompts or use a similar public dataset.
The main limitation is the reliance on a specific TTS backbone (F5-TTS), which may not generalize to other architectures without significant modification. The accent control is limited to the accents seen during training or those that can be represented by the exemplar encoder; truly novel accents not in the training distribution may not be well-captured. The listening test is small-scale, and the objective accent probe may not fully capture perceptual accent nuances. The method requires separate exemplar clips for accent, which adds a data requirement at inference time.
The work has significant implications for personalized TTS systems, allowing users to control not just the voice but also the accent of the synthesized speech. This could be useful for language learning, accessibility, and creative applications. The decoupling of identity and accent is a step towards more fine-grained control over speech synthesis, which is a key goal in the field. The method's ability to generalize to out-of-domain accents is particularly promising for real-world applications where the target accent may not be in the training set. The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Primary: Unknown
All Institutions: Unknown
The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
The paper proposes a principled two-dimensional taxonomy for spoken conversational memory, crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal). This is a significant methodological advance over prior benchmarks that treated memory as a single-session, lexical-only problem. The construction pipeline is rigorous, utilizing a three-stage process (planning, dialogue writing, speech synthesis) with specific controls to ensure "acoustic necessity" (i.e., the answer cannot be derived from the transcript alone). The use of Higgs-TTS-3 and VCTK voices, along with mixed environmental sounds from ESC-50, provides a controlled yet realistic audio environment. The inclusion of "haystack" and "filler" sessions to create multi-session histories of varying lengths (8K-64K tokens) is a strong design choice that isolates the effect of context length from question difficulty.
The evaluation covers 15 Large Audio Language Models (LALMs), including both open-weight and proprietary models. The results reveal a clear hierarchy of difficulty: models perform significantly better on speech semantics than on audio-native cues (speaker, paralinguistic, environmental). The finding that no model exceeds 40% accuracy at the 32K token budget highlights a critical gap in current LALM capabilities. The error analysis is particularly valuable, distinguishing between binding failures (wrong speaker) and retention failures (lost cue), which provides actionable insights for future model development. The controlled scaling analysis shows that performance degrades as history length increases, with non-lexical information degrading faster than lexical information.
The paper provides detailed descriptions of the construction pipeline, quality control checks, and evaluation protocol. The use of specific TTS systems, voice datasets, and sound event datasets enhances reproducibility. The release of the benchmark instances, question templates, and judge prompts (mentioned in the text) further supports reproducibility. However, the reliance on proprietary LLMs (Gemini-3.7-Flash, GPT-5.6-Luna) for generation and evaluation introduces a dependency on external services, which may limit full reproducibility for all researchers.
The benchmark relies on synthetic speech generated by TTS systems, which may not fully capture the variability and noise of real-world human speech. The "acoustic necessity" check, while rigorous, uses a specific LLM (Gemini-3.7-Flash) for validation, which could introduce bias if that model has specific strengths or weaknesses in audio understanding. The paper does not extensively discuss the computational cost of evaluating models on 64K token audio histories, which may be prohibitive for some researchers. Additionally, the focus on English speech (implied by the datasets used) limits the generalizability of the findings to other languages.
VoxMem addresses a fundamental challenge in the development of long-term conversational AI systems. By providing a standardized benchmark for multi-session, audio-native memory, it enables fair comparison of LALMs and guides future research toward improving non-lexical acoustic memory. The taxonomy and benchmark are likely to become a standard reference in the field, driving progress in speaker identification, paralinguistic understanding, and environmental sound recognition within conversational contexts. The findings on the distinct failure modes for different evidence types will inform the design of specialized memory modules in LALMs. The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Primary: KAIST
All Institutions: KAIST, VGG, University of Oxford
The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
The paper employs a rigorous mechanistic interpretability framework, combining Representational Similarity Analysis (RSA) and Causal Mediation Analysis (CMA) to dissect the internal workings of Audio-Visual LLMs. The authors successfully identify a three-stage symbolic binding mechanism (Anchor ID Retrieval, Target ID Selection, Feature Retrieval) that relies on modality-specific symbolic variables (Temporal IDs for audio, Position IDs for vision). This is a sophisticated approach that moves beyond black-box evaluation to understand *how* models process cross-modal information. The proposed intervention, using an off-the-shelf Active Speaker Detection (ASD) model to overlay bounding boxes, is a clever, training-free (and lightly fine-tuned) solution that directly targets the identified bottleneck in audio-visual alignment.
The experiments are extensive and well-structured. The authors validate their mechanistic findings across four different AVLLM architectures (video-SALMONN2+, Qwen2.5-Omni, MiniCPM-o-4.5) using both synthetic toy datasets and real-world benchmarks (SocialOmni, AVSpeaker, DailyOmni, etc.). The ablation studies effectively isolate the contribution of the ASD prompting, showing that it outperforms other training-free decoding methods and standard fine-tuning. The generalization to broader audio-visual benchmarks (DAVE, OmniBench, WorldSense) further strengthens the claim of the method's utility.
The paper provides detailed implementation specifics, including LoRA ranks, training steps, and dataset construction methods. The use of standard off-the-shelf models and clear descriptions of the prompting strategy enhances reproducibility. However, the specific synthetic video generation pipeline and the exact ASD model used (referenced as [CITATION]) would need to be clearly specified in the final publication for full reproducibility.
The primary limitation is the reliance on a synthetic toy dataset for the core mechanistic analysis, which may not fully capture the complexity of real-world multi-speaker scenarios, although the authors do validate on real-world data. Additionally, the method depends on the accuracy of the external ASD model; if the ASD model fails, the prompting strategy may degrade performance. The fine-tuning, while lightweight, still requires access to the model's weights and a GPU, which may not be feasible for all users.
This work has significant implications for the development of more robust multimodal AI systems. By identifying specific failure modes in cross-modal binding, it provides actionable insights for improving model architectures and training strategies. The proposed ASD prompting method is a practical, low-cost solution that can be easily integrated into existing AVLLM pipelines, potentially improving performance in applications like video conferencing, accessibility tools, and video understanding. The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
Primary: Sakana AI
All Institutions: Sakana AI
The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
The paper proposes a "symbiotic" architecture where an audio injector module generates the Key-Value (KV) cache for a frozen Large Language Model (LLM), bypassing the need to pass audio embeddings through the LLM's self-attention layers during prefilling. The injector is a CNN-based module (using LConv blocks) that maps audio encoder outputs directly to the LLM's KV space. Key technical contributions include a "KV Scale Matching" strategy to align the injector's output distribution with the LLM's internal KV distribution, and "Noisy RoPE" training to improve robustness to sequence lengths unseen during training. The approach is theoretically sound, leveraging the fact that the KV cache is the interface between context and generation, allowing modality-specific processing to be decoupled from the backbone.
Experiments are conducted on a compact model (Qwen3-0.6B) with WavLM as the audio encoder. The paper evaluates ASR (LibriSpeech), Audio QA (Clotho), and Acoustic Scene Classification (CochlScene). The proposed method outperforms the "Encoder-only" (SLM-style) baseline significantly on non-ASR tasks and approaches the performance of a fine-tuned "Monolithic" model while using fewer active parameters during audio prefilling. Crucially, it preserves text-only performance (WikiText-2, HellaSwag, GSM8K) by construction, avoiding catastrophic forgetting. However, the evaluation is limited to a single small backbone, and the speedup is modest (157s vs 199s) in the tested regime.
The paper provides detailed architectural descriptions, including specific hyperparameters for the injector (kernel size, subsampling stride), training configurations (learning rates, batch size, steps), and specific techniques for stabilization (RMSNorm initialization, noisy RoPE parameters). The use of standard open-source components (WavLM, Qwen3) enhances reproducibility. However, code availability is not explicitly confirmed in the text, and the specific implementation of the KV injection into the inference engine is not detailed.
The primary limitation is the scale of the experiments; using a 0.6B LLM does not fully validate the scalability claims for larger models where the prefilling bottleneck is more severe. The speedup demonstrated is relatively small in the current setup, and the authors acknowledge that the CNN-centric injector may limit performance on complex acoustic tasks compared to attention-based injectors. The method relies on specific internal details of the backbone (e.g., Qwen3's RMSNorm), which may require adaptation for other LLM families.
This work offers a practical solution to the high computational cost of multimodal prefilling and the risk of catastrophic forgetting in LLM fine-tuning. By decoupling audio processing from the backbone, it enables the use of larger, more capable LLMs for audio tasks without proportional increases in inference latency for the audio portion. This could facilitate the deployment of ALMs in edge devices or real-time applications where prefilling latency is critical. The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.
Primary: Sungkyunkwan University
All Institutions: Sungkyunkwan University
The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
The paper proposes a novel "acoustic-to-text KV compression" strategy for full-duplex speech models. The core idea is to exploit "listening-time slack"โthe computational gap between audio unit arrivalsโto run a lightweight transcription side channel (implemented via LoRA) that converts incoming speech into text tokens. These text tokens are retained in the KV cache while older acoustic KV states are evicted. The methodology is technically sound, leveraging the existing language model backbone rather than a separate ASR model. A significant methodological strength is the use of knowledge distillation from the frozen original model to preserve native listening/speaking behaviors (turn-taking, interruption) which would otherwise be disrupted by the new transcription objective. The training setup uses word-level forced alignments and a composite loss function including cross-entropy for ASR and KL divergence for behavior preservation.
The experiments are conducted on the MiniCPM-o 4.5 model. The evaluation covers three key areas: 1) Long-form speech understanding (LongSpeech benchmark), showing improved WER and QA accuracy compared to native streaming, and competitive performance with external ASR cascades. 2) Full-duplex interaction (Full-Duplex-Bench), demonstrating that the distillation component is critical for maintaining natural conversational dynamics (pause handling, turn-taking). 3) Efficiency, showing a 64.6% reduction in peak KV cache size with no real-time deadline misses on H100 GPUs. The ablation study effectively isolates the impact of the distillation loss, showing that without it, the model fails at turn-taking. The comparison with external ASR cascades is particularly relevant, showing that the proposed method achieves similar quality with fewer added parameters and integrated memory management.
The paper provides sufficient detail for reproduction, including the specific model (MiniCPM-o 4.5), training data (LibriSpeech train-clean), hyperparameters (LoRA rank 16, learning rate, batch size), and the specific eviction strategy (5-unit retention window). The use of standard benchmarks (LongSpeech, Full-Duplex-Bench) facilitates comparison. However, the code is not explicitly linked in the text provided, and the specific implementation details of the "listening-time slack" scheduling might require access to the source code for precise replication.
The method is evaluated primarily on a single model architecture (MiniCPM-o 4.5), so generalizability to other full-duplex models is not established. The transcription side channel relies on the model's ability to predict the start of the ASR segment; errors in this gating mechanism could lead to missed transcriptions. The method assumes a consistent "listening-time slack" of ~900ms, which may not hold for all hardware configurations or model variants. Additionally, the WER (12.9%) is higher than dedicated streaming ASR models (e.g., FastConformer at 10.3%), suggesting a trade-off between integration and pure transcription accuracy.
This work addresses a critical bottleneck in deploying full-duplex speech agents: memory consumption during long conversations. By converting acoustic history to text, it enables longer context windows without proportional memory increases. This is highly relevant for real-time voice assistants and interactive agents. The approach of using "slack" time for auxiliary tasks (transcription) is a generalizable principle for efficient inference in streaming multimodal models. The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Primary: Microsoft
All Institutions: Microsoft, University of Michigan
The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
The paper proposes a pragmatic and effective adaptation of a deployed personalized speech enhancement model (PVQE) into an online audio-visual target-speaker extraction system (V+E). The core methodological contribution is the integration of visual cues (mouth motion via AV-HuBERT) into the existing speaker-conditioning path of the enhancer, rather than training a separator from scratch. The use of a gated residual connection to fuse enrollment embeddings with visual features is a sound architectural choice that allows the model to retain its enhancement capabilities while learning selection. The fine-tuning strategy, which involves removing lookahead operations and rearranging decoder filters for causal processing, is well-described and technically rigorous. The approach effectively leverages the prior knowledge of the enhancement model to solve the extraction problem, addressing the "selection" gap in personalized enhancement.
The experimental evaluation is extensive and high-quality. The authors test on both synthetic mixtures (LRS3, VoxCeleb2) and, crucially, recorded meeting corpora (AMI, MCoRec, MTM, UniTalk), which is a significant strength as it moves beyond the standard synthetic benchmark limitations. The inclusion of a large-scale human listening test (170 listeners, P.835 protocol) provides strong evidence for the perceptual quality claims. The results show consistent improvements over the baseline extractor (AVASE) in both objective metrics (SI-SNRi, WAcc) and subjective quality (MOS). The preservation and rejection tests further validate the system's deployability, showing it does not distort the target when no interference is present. The analysis of the UniTalk results, where background noise suppression degrades, is honest and provides valuable insight into the trade-offs of this adaptation strategy.
The paper provides sufficient detail for reproducibility, including model architecture dimensions, training data sources, and fine-tuning procedures. The use of standard public datasets (LRS3, VoxCeleb2, AMI, MCoRec) aids reproducibility, although the internal MTM dataset is not publicly available. The specific hyperparameters and initialization strategies (e.g., zero-initialization of the enrollment projection) are clearly stated. However, the code and model weights are not explicitly linked in the provided text, which is a minor drawback for immediate reproducibility.
The primary limitation is the degradation in background noise suppression on in-the-wild data (UniTalk) compared to the original enhancement model, indicating that the adaptation to extraction comes at a cost to general noise robustness. The system relies on the presence of synchronized video, which may not be available in all deployment scenarios. Additionally, the comparison with PVQE (the starting model) is somewhat confounded because PVQE is not designed for extraction, so the "quality retention" claim is relative to a model that fails at the primary task (selection) in many cases.
This work has significant practical impact for real-time communication systems (e.g., video conferencing). By demonstrating that a deployed enhancement model can be adapted for extraction with minimal architectural changes and high perceptual quality, it offers a viable path for improving user experience in noisy or multi-speaker environments. The focus on low-latency, online processing makes it highly relevant for industrial applications. The rigorous evaluation on recorded meetings sets a new standard for how such systems should be tested, encouraging the field to move beyond synthetic benchmarks. The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
Primary: Purdue University
All Institutions: Purdue University, Loyola University Chicago, University of Michigan
The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
The paper proposes NoteSep, a two-stage framework for score-informed note separation. The first stage, NoteGrab, utilizes a dual-stream TFC-TDF U-Net architecture to process harmonic and percussive components (via HPSS) separately, linked by bidirectional cross-attention. A key methodological contribution is the "selective harmonic gating," which conditionally suppresses lower-octave interference based on score alignment, addressing a specific physical challenge in polyphonic music. The second stage, Adaptive Set Ownership (ASO), is a novel joint separation module that reallocates mixture energy among concurrent note estimates to enforce consistency, using a learned gate and log-gain prediction. This approach effectively tackles the "double counting" problem inherent in independent note extraction.
The authors curate a substantial training set (SCNS-Train) and a disjoint evaluation set (SCNS-Eval) covering 16 instruments. The evaluation is rigorous, comparing against both traditional methods (Score-Informed NMF) and commercial software (Melodyne). The results show a significant improvement in SI-SDR (7.39 dB vs 2.49 dB for the strongest baseline). Furthermore, the paper evaluates the system on real-world ensemble recordings (PHENICX-Anechoic, Bach10) by aggregating note estimates into instrument stems, demonstrating competitive performance against specialized instrument separation models. The inclusion of an editing evaluation (SCNS-Edit) further validates the utility of the separated notes for downstream tasks.
The paper commits to releasing code, model weights, and datasets. The architectural details are described with sufficient specificity (e.g., TFC-TDF v3, FiLM conditioning, specific loss functions), and the demo page provides audio examples. The use of standard libraries (librosa) for preprocessing aids reproducibility.
The method relies heavily on accurate score alignment (pitch, onset, offset); errors in the input score will propagate to the separation quality. The computational cost is non-trivial, requiring K passes for K notes plus an ASO pass, which may limit real-time application on consumer hardware. The evaluation on real recordings is indirect (via instrument stem aggregation), so the fidelity of individual notes in complex, noisy real-world scenarios is not directly quantified with the same rigor as the synthetic benchmark.
This work enables precise, note-level audio editing, which has significant applications in music production, education (isolating specific notes for learning), and restoration. It bridges the gap between symbolic music information and audio signal processing, potentially facilitating more granular music information retrieval tasks. The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
Primary: University of Pisa
All Institutions: University of Pisa
The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
The paper employs a rigorous three-step interpretability pipeline to analyze Part-of-Speech (PoS) encoding in Sparse Autoencoder (SAE) latents. First, it establishes recoverability using L1-regularized logistic regression probes on SAE activations from LLaMA-3-8B. Second, it moves beyond binary accuracy to localize information by analyzing feature salience (classifier coefficients) and coverage (the minimal number of latents required to cover 95% of instances for a given PoS tag). Third, it validates these localized groups on held-out data and a controlled synthetic dataset to test stability and additivity. The methodology is sound, moving logically from "is the information there?" to "where is it?" and "is it stable?". The use of a controlled dataset with minimal templates is a strong methodological choice to isolate lexical effects from syntactic ones.
The experiments are comprehensive, utilizing the GUM Treebank for natural data and a custom controlled dataset for validation. The results clearly demonstrate that PoS information is distributed across compact groups of latents rather than single monosemantic units. The distinction between Open-class (nouns, verbs) and Closed-class (determiners, conjunctions) PoS is well-supported, with closed classes showing more compact and stable latent groups. The ablation studies, including random label controls and layer-wise analysis, effectively rule out trivial explanations like lexical memorization or layer-specific artifacts. The finding that a small union of these latents (498 out of 131,072) supports strong multi-class classification is a significant quantitative result.
The authors provide high reproducibility by releasing code, data, and the specific SAE checkpoint used. The GitHub repository contains the pipeline for extracting activations and running the probes. The use of standard libraries (Scikit-learn, Sparsify) and public models (LLaMA-3-8B, EleutherAI SAE) ensures that other researchers can easily replicate the findings. The detailed description of the subword-to-token alignment strategy (leftmost subword anchoring) is crucial for reproducibility in this domain.
The study is limited to a single model (LLaMA-3-8B) and a single SAE variant, which may limit the generalizability of the findings to other architectures or sparsity regimes. The analysis is restricted to English, and it is unclear how these patterns would manifest in morphologically rich languages. The controlled dataset, while useful, is small and covers only a subset of syntactic structures. Additionally, the reliance on linear probes means that non-linear interactions between latents are not captured, potentially underestimating the complexity of the representation.
This work contributes to the broader field of mechanistic interpretability by providing a structured framework for analyzing how linguistic categories are represented in sparse latent spaces. It challenges the assumption of strict monosemanticity for linguistic features, suggesting instead a distributed but localized organization. This has implications for how interpretability tools are designed and how we understand the internal representations of LLMs. The findings could inform the development of more targeted interpretability methods that focus on groups of features rather than individual units. The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Primary: Xiamen University
All Institutions: Xiamen University
[One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.
The paper proposes EditVoice, a non-autoregressive (NAR) zero-shot TTS model based on Edit Flows. The core methodological contribution is the application of Edit Flows to speech token sequences, enabling variable-length generation through insertions, deletions, and substitutions without pre-specifying sequence length. The authors introduce a speech-infilling training objective that unifies TTS and speech editing. A key heuristic contribution is Complementary Prompt Sampling (CPS), which leverages the complementary predictions of prefix and suffix prompt placements to improve generation quality, along with a local consistency rule to prevent artifacts. The method also exploits the model's ability to generalize to editing plausible (non-noise) token sequences for post-generation refinement.
Experiments are conducted on Seed-TTS Eval EN and LibriSpeech-PC for TTS, and RealEdit for speech editing. The model is trained on 10K hours of GigaSpeech. Results show competitive WER and speaker similarity compared to autoregressive baselines like CosyVoice2 and other NAR models. The 16-NFE variant achieves a low RTF of 0.0989. Ablations confirm the effectiveness of CPS and post-generation refinement. Subjective evaluations (NMOS, SMOS) are included.
The paper provides detailed architectural specifications (14-layer LLaMA-style Transformer, Conformer text encoder) and training hyperparameters. It specifies the use of frozen S3Tokenizer2 and CosyVoice2 acoustic decoder. However, the code repository is not explicitly linked in the text provided (only a demo page), which may limit immediate reproducibility compared to papers with open-source code.
The model relies on a specific tokenizer (S3Tokenizer2) and acoustic decoder (CosyVoice2), which may limit generalizability to other speech tokenization schemes. The CPS heuristic, while effective, is a manual rule-based approach that may not scale optimally to all conditions. The training data size (10K hours) is significantly smaller than some state-of-the-art systems (1M+ hours), though the paper argues for efficiency.
This work contributes to the efficiency of NAR TTS by removing the need for length prediction, a common bottleneck. The unification of TTS and editing in a single model is a step towards more flexible speech manipulation tools. The use of Edit Flows for speech is a novel application of this generative framework. [One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.