Total: 1
A human face conveys rich cues about speaker identity, enabling face-based zero-shot text-to-speech (TTS) for unseen speakers. However, in modular face-based TTS systems, the acoustic model is typically trained on speech-derived embeddings, while face-derived representations are introduced only at inference time, often resulting in identity drift. We propose Dual-Space Constrained TTS (DSC-TTS), a modular framework that enforces identity consistency during acoustic model training in both the speaker embedding space and a shared identity space learned through face-voice alignment. By constraining representations across these complementary spaces, the proposed framework improves speaker identity stability while preserving speech quality. Experiments demonstrate higher speaker similarity and stronger identity consistency than existing face-based TTS methods.