Total: 1
Text-to-image models ignore paralinguistic cues. Pitch, rate, and emotional inflection shape how listeners visualize a scene, yet audio-to-image methods treat speech as a single opaque vector. NovaDiffusion conditions diffusion-based synthesis on emotion-correlated prosodic features extracted alongside general audio representations. It comprises (1) Prosody-CLIP, trained on 168K human-annotated speech-emotion-image triplets to align prosody-enriched embeddings with visual representations; (2) a 280M-parameter distilled U-Net with multi-scale fusion; and (3) a decoupled cross-attention adapter injecting prosodic embeddings via IP-Adapter. On RAVDESS (8 emotions, chance 12.5%), NovaDiffusion achieves 71.3% ECA, outperforming SonicDiffusion by 23.1pp; on held-out IEMOCAP speakers it reaches 63.4% vs. 41.7%. ECA relies on facial expression classification and does not capture scene-level emotion. Scope is limited to emotion-correlated prosody; full prosodic control remains open.