| Total: 177
Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling and fine-grained control for coherent multicast audiobooks. To address this, we propose a context-aware, emotion-controllable speech synthesis framework with three innovations: a context mechanism for consistency, a disentanglement paradigm to decouple style from prompts, and self-distillation to boost expressiveness. Experiments show consistent gains: chapter-level generation achieves 4.25 M-MOS(14% relative gain over the strongest baseline), dialog reaches 4.11 S-MOS, and emotion control improves by 31 percentage points in high-intensity discrimination. Ablation studies validate our methods. Demo: https://semisemi-ux.github.io/.
Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.
Despite progress in text and visual generation, coherent long-form audio storytelling remains challenging. Existing systems often suffer from mismatches between character settings and voice performance, weak self-correction, and limited user interaction. We propose AuDirector, a self-reflective closed-loop multi-agent framework for audio narrative generation. Its identity-aware pre-production mechanism converts narratives into character profiles and utterance-level emotion instructions, retrieves suitable voices, and guides expressive speech synthesis. A collaborative synthesis and correction module audits and regenerates defective audio via closed-loop self-correction. A human-guided interactive refinement module further interprets natural language feedback to revise scripts interactively. Experiments show that AuDirector outperforms state-of-the-art baselines in structural coherence, emotional expressiveness, and acoustic fidelity. Samples: https://github.com/Riddae/AuDirector.
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations, such as monster growls and robotic voices, underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, constructed by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to produce corresponding effect-modified variants. We further provide a standardized test set with explicit seen/unseen splits over source timbre groups and preset styles to assess generalization under controlled conditions. Finally, we report baseline benchmark results to support reproducible evaluation and future research. The dataset and demo samples are available online.
Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. We propose ZipL-Dialog, which shifts conditional flow-matching into a 4x time-compressed (25 Hz) latent space. To preserve acoustic fidelity under compression, we employ a deterministic mel autoencoder with auxiliary mel-domain supervision and optimize the ZipFormer's hierarchical downsampling schedule. Experiments show that ZipL-Dialog reduces maximum peak GPU memory by 11.22x and accelerates inference by 2.23x over the baseline, substantially lowering the memory footprint of single-pass multi-minute dialog synthesis while maintaining perceptual naturalness. Audio samples are available at https://speechdemos.github.io/.
Standard TTS metrics such as MOS and mel-cepstral distortion provide global scores but do not locate where synthesis diverges from natural speech. To examine this, four systems (Tacotron2-DDC, FastSpeech2, Glow-TTS, MixerTTS) are analysed across 13,100 matched LJ-TTS utterances. At the prosodic level, global F0 variability is compressed (d = −0.55) while local pitch reversals increase (d = +0.82), suggesting a cross-timescale dissociation, rather than monotone intonation. Timing metrics show no reliable deviation, localizing the deficit to F0 coordination. At the segmental level, vowel spaces shrink to 9–30% of the human baseline, directional formant biases indicate articulatory undershoot, and locus analysis confirms reduced place-conditioned coarticulation most consistently for alveolars. Deviations at these two levels are largely uncorrelated, suggesting prosodic organisation and segmental precision are distinct dimensions of synthesis quality conflated by standard metrics.
Achieving fine-grained emphasis control in speech synthesis remains challenging due to data scarcity and the inherent complexity of prosody. To address this issue, we extend F5-TTS with an additional Emphasis Encoder and propose a three-stage optimization framework that progressively enhances emphasis controllability. The framework begins with Supervised Fine-Tuning on manually annotated data, followed by Direct Preference Optimization, where preference pairs are constructed by ranking SFT-generated samples using the Wavelet Prosody Toolkit (WPT). In the final stage, we adapt Flow-CPS to perform online reinforcement learning (RL) for flow-matching models. Using WPT as a reward model, this stage refines the flow trajectories through the estimation of the relative advantage of the group. Experimental results demonstrate that the proposed pipeline substantially improves emphasis intensity and controllability while preserving natural prosody. Audio samples are available at https://thuhcsi.github.io/interspeech2026-F5Emphasis.
While zero-shot text-to-speech models perform well in global voice cloning, they struggle with fine-grained prosodic control, as strict word-level intensity and tempo manipulation can disrupt acoustic priors and introduce artifacts. We propose CraftTTS, a three-stage framework enabling stable word-level control without sacrificing fluency. First, a compute-driven zero-shot pipeline automatically constructs large-scale preference pairs without manual annotation. Second, joint supervised fine-tuning and direct preference optimization improve sensitivity to local prosodic tags. Third, group relative policy optimization with a multi-dimensional prosodic reward balances local controllability and global naturalness by regulating intelligibility, intensity contrast, and rhythm. Experiments demonstrate that CraftTTS achieves state-of-the-art fine-grained expressiveness while preserving zero-shot generation capability.
Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based TTS methods ignore the style-specific prosodic patterns in generated speech, resulting in deficient style learning and thus limiting speaker similarity in synthesized speech. To this end, we investigate the prosody learning conditioned on the synthesized speech, and propose to predict the prosody of the current syllable based on previously predicted speech. Experimental results obtained on three datasets demonstrated the efficacy of the proposed dynamic prosody prediction method in enhancing the prosody learning capability, thereby improving the speaker similarity of the generated speech. Audio samples are available at https://muzw.github.io/dynapros/.
Non-autoregressive (NAR) text-to-speech (TTS) models excel in parallel inference and style consistency. However, their reliance on accurate character-level or global duration predictions remains a critical bottleneck, limiting both synthesis fidelity and naturalness. This paper proposes ElasticDLM (Elastic-length Diffusion Language Model), a novel variable-length NAR TTS model that eliminates the strong dependency on accurate duration predictions. We introduce two functional tokens, along with a Differentiated Length-Scaling Training Scheme (DLTS) and a Hierarchical Confidence-Guided Inference (HCGI) strategy tailored for speech characteristics, enabling the model to learn length-scaling features and automatically adjust sequence length during inference. Experiments show that ElasticDLM produces high-fidelity, natural speech from arbitrary-length inputs, demonstrating a more flexible and simplified approach to high-quality speech synthesis. Audio samples are available at https://thuhcsi.github.io/ElasticDLM/.
We present OmniVoice, a massively multilingual zero-shot text-to-speech (TTS) model that scales to over 600 languages. At its core is a novel diffusion language model-style discrete non-autoregressive (NAR) architecture. Unlike conventional discrete NAR models that suffer from performance bottlenecks in complex two-stage (text-to-semantic-to-acoustic) pipelines, OmniVoice directly maps text to multi-codebook acoustic tokens. This simplified approach is facilitated by two key technical innovations: (1) a full-codebook random masking strategy for efficient training, and (2) initialization from a pre-trained LLM to ensure superior intelligibility. By leveraging a 581k-hour multilingual dataset curated entirely from open-source data, OmniVoice achieves the broadest language coverage to date and delivers state-of-the-art performance across Chinese, English, and diverse multilingual benchmarks. Our code and pre-trained models are publicly available.
Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention. In this paper, we propose WAND, Windowed Attention and Knowledge Distillation, a framework that adapts pretrained AR-TTS models to operate with constant computational and memory complexity. WAND separates the attention mechanism into two: persistent global attention over conditioning tokens and local sliding-window attention over generated tokens. To stabilize fine-tuning, we employ a curriculum learning strategy that progressively tightens the attention window. We further utilize knowledge distillation from a full-attention teacher to recover high-fidelity synthesis quality with high data efficiency. Evaluated on three modern AR-TTS models, WAND preserves the original quality while achieving up to 66.2% KV cache memory reduction and length-invariant, near-constant per-step latency.
Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https: //voicetta.pages.dev/.
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality. Audio samples are available on our web page1.
A human face conveys rich cues about speaker identity, enabling face-based zero-shot text-to-speech (TTS) for unseen speakers. However, in modular face-based TTS systems, the acoustic model is typically trained on speech-derived embeddings, while face-derived representations are introduced only at inference time, often resulting in identity drift. We propose Dual-Space Constrained TTS (DSC-TTS), a modular framework that enforces identity consistency during acoustic model training in both the speaker embedding space and a shared identity space learned through face-voice alignment. By constraining representations across these complementary spaces, the proposed framework improves speaker identity stability while preserving speech quality. Experiments demonstrate higher speaker similarity and stronger identity consistency than existing face-based TTS methods.
Most existing speech enhancement operate directly in the acoustic space, where noise and speech components are inherently entangled, leading to artifacts such as high-frequency attenuation and background holes. In this paper, we propose Seed-Enh, a generative speech enhancement framework that performs enhancement in decoupled semantic and timbre spaces. Our framework processes noisy speech in three stages: (1) semantic space processing using a frozen Whisper encoder to extract noise-robust semantic representations, (2) timbre space processing using CAM++ embeddings and context learning, and (3) space fusion via flow matching. Experiments demonstrate that Seed-Enh achieves the highest overall quality scores compared to state-of-the-art baselines, effectively suppressing noise while restoring harmonic structures and high-frequency details. Beyond speech enhancement, Seed-Enh also enables zero-shot voice conversion, significantly outperforming Seed-VC when handling noisy inputs. The demo and pretrained model are available at https://github.com/shangqwe123/Seed-Enh.
Noise-robust bandwidth expansion aims to reconstruct high-fidelity wideband speech from noisy low-resolution inputs. While flow matching has shown strong performance in speech generation, accurately recovering clean speech from noisy inputs remains challenging due to the ambiguity of velocity estimation under noise. In this work, we propose VeRe-Flow, a clean-guided flow matching framework that introduces multi-level clean supervision to guide the generative process toward clean speech. At the velocity level, we introduce velocity contrastive regularization, which attracts the predicted velocity toward the clean trajectory while repelling it from noisy trajectories. At the representation level, we incorporate representation alignment that aligns intermediate features with clean self-supervised learning representations. The results demonstrate that the proposed method achieves the lowest LSD and highest DNSMOS OVRL among all baselines, and the highest MOS among generative baselines.
Generative speech enhancement methods have shown impressive performance by directly modeling clean speech distribution, yet their effectiveness critically depends on the reliability of conditional information. Mainstream conditioning strategy face two primary challenges: the shallow features extracted from noisy speech often fails to capture structured acoustic features such as harmonic structures, and extracting accurate and reliable semantic information from noisy signals poses a challenge comparable to the enhancement task itself. To address these limitations, we propose HFMSE, a Harmonic-guided Flow Matching method for Speech Enhancement. Specifically, we design an efficient harmonic encoder that extracts harmonic structural features from noisy speech through a two-step process of fundamental frequency localization and harmonic mask generation. These features serve as strong conditioning guidance and are integrated into the flow matching generation process to improve the structural integrity and perceptual naturalness of the output speech. Extensive experiments on DNS Challenge 2020 show that HFMSE achieves state-of-the-art performance, demonstrating its superior capability and robustness. Code is available at https://github.com/xxnhq/HFSE/.
Flow matching (FM) enables high-fidelity generation, while self-supervised learning (SSL) speech models provide hierarchical representations spanning acoustic and phonetic levels. However, existing FM-based speech enhancement (SE) methods operate primarily in the spectral domain, treating SSL features only as external conditions rather than modeling directly in the SSL latent space. To fully exploit the structural richness of SSL representations, we propose PhASE-Flow, an FM-based SE framework that operates entirely in the SSL space. It models the conditional distribution of clean acoustic representations given phonetic ones, reconstructing the waveform via a neural vocoder. Experiments show that PhASE-Flow outperforms state-of-the-art baselines in perceptual quality and intelligibility. Notably, it achieves competitive performance with only four sampling steps, enabling highly efficient inference. Audio demos are available at https://anonymous.4open.science/w/phase-flow_demo-E6E1/.
Diffusion models show potential for speech enhancement but lack linguistic guidance. We condition a diffusion-based model on wav2vec 2.0 features from noisy input, injected at the U-Net bottleneck via Feature-wise Linear Modulation (FiLM). Phonetic representations from wav2vec 2.0 features of degraded speech, anchor the reverse diffusion process. While a frozen wav2vec 2.0 encoder extracts features, a learned FiLM generator produces scale and shift parameters modulating the bottleneck with minimal overhead. Motivated by the optimal Bayesian causal estimator under a linear-Gaussian state-space model, FiLM coefficients are aggregated via exponential smoothing for temporal compression. Evaluation on VoiceBank-DEMAND and LibriMix shows competitive performance against the unconditioned baseline in PESQ, STOI, SI-SDR and DNSMOS. We consistently record an improvement of 0.4 on PESQ score, suggesting self-supervised representations effectively condition diffusion-based speech enhancement.
Deep learning-based speech enhancement (SE) models are typically trained on synthetic pairs of clean speech and synthesized noisy speech, which often generalize poorly to real-world acoustic environments. In practical target environments, however, paired clean target speech aligned with noisy recordings is usually unavailable, making it difficult to directly use real data for supervised SE training. To address these limitations, we propose SwitchSE — a target-domain clean-free fine-tuning framework with a switch-controlled mechanism that leverages transcription-noisy speech pairs to upgrade a pre-trained SE model in realistic environments. Experiments show that with only 2.9 h of CHiME-3 real noisy speech, SwitchSE substantially improves target-domain performance while retaining strong performance on the original synthetic dataset, providing a data-efficient framework to bridge the gap between synthesized training data and real-world application demands.
Using speaker embeddings as conditioning can strengthen speech enhancement, but most methods either require clean enrollment audio or rely on embeddings extracted from noisy speech, which are fragile under noise and domain shift. We propose G-MaP-SE, a guided enhancement framework that builds a clean-speech embedding prior with a Gaussian Mixture Model (GMM) and refines a noisy conditioning embedding by matching it to this prior. The matched prior embedding is then injected into a time-frequency enhancement backbone via a lightweight gated fusion module. Experiments on VoiceBank+DEMAND and DNS Challenge 2020 datasets show that the proposed prior matching consistently outperforms noisy conditioning and substantially narrows the gap to an oracle clean-conditioning upper bound, while requiring no enrollment audio at inference time. The code, audio samples, and checkpoint are available.
Deep learning-based speech enhancement typically relies on paired data, creating a domain gap between training and deployment. Unsupervised methods based on generative adversarial networks (GANs) use unpaired data as source priors, but existing single-branch models often suffer from source leakage due to a dominant consistency loss or a weak clean speech prior. In this work, we present a comprehensive analysis of prior data in unsupervised GAN-based speech enhancement and utilize a dual-branch framework that explicitly models both clean speech and noise. Our experiments show that mismatched priors increase source leakage. By utilizing aligned noise prior data—which are easy to collect in practice—we significantly mitigate speech-noise leakage, prevent the over-suppression of speech, and improve perceptual quality. This demonstrates that leveraging environment-specific noise data is an effective strategy for improving unsupervised enhancement in target domains.
Most deep learning based speech enhancement methods are usually trained in a supervised fashion, i.e., they typically rely on parallel corpora of noisy and clean speech pairs. This is often difficult to obtain in real-world scenarios, leading to the use of synthetic data. In this work, we propose a novel unsupervised speech enhancement method that does not require paired training data. We introduce a multi-discriminator GAN-based architecture to capture global (or utterance-level) and local (or frame-level) characteristics of the speech signal. Additionally, we incorporate self-supervised representations from a pre-trained model to provide auxiliary information to the generator and thereby enhance the denoising capability. Our extensive experimental results on the VoiceBank+DEMAND dataset demonstrate that the proposed method achieves comparable or better performance across both intrusive and non-intrusive quality measures.
Continuous-time diffusion models have demonstrated strong capabilities for modeling complex data distributions across various domains. However, these models often encounter gradient conflicts, where parameter updates at different timesteps interfere with each other, hindering effective training. To provide theoretical insights into this phenomenon, we introduce Delta Loss Lower Bound (DELLBO), a tractable lower bound on the reduction in the integrated training loss. Through mathematical analysis of DELLBO, we derive the optimal parameter update direction as the opposite of the vector from the origin to the nearest point in the convex hull of per-timestep gradient vectors. Building on this insight, we propose Self-adaptive Gradient Conflict Mitigator (SGCM), a simple yet effective method that adjusts parameter update directions during training. The results confirm that SGCM behaves consistently with our theoretical analysis and demonstrate its effectiveness for speech enhancement.