LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

#1 LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models [PDF²] [Copy] [Kimi¹] [REL]

Authors: Danilo de Oliveira, Julius Richter, Tal Peer, Timo Germann

We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 and TCD-TIMIT demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition (ASR). These findings are also supported by a formal listening experiment. Extensive ablation studies and cross-dataset evaluation confirm the effectiveness and generalization capabilities of our approach.

Subjects: Audio and Speech Processing , Sound

Publish: 2025-05-16 15:56:07 UTC

2505.11391

#1 LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models [PDF2] [Copy] [Kimi1] [REL]

#1 LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models [PDF²] [Copy] [Kimi¹] [REL]