tiwari26@interspeech_2026@ISCA

Total: 1

#1 Say That Again: Visualizing Paralinguistic Cues with Prosody-Aware Diffusion [PDF] [Copy] [Kimi] [REL]

Author: Shyamji Tiwari

Text-to-image models ignore paralinguistic cues. Pitch, rate, and emotional inflection shape how listeners visualize a scene, yet audio-to-image methods treat speech as a single opaque vector. NovaDiffusion conditions diffusion-based synthesis on emotion-correlated prosodic features extracted alongside general audio representations. It comprises (1) Prosody-CLIP, trained on 168K human-annotated speech-emotion-image triplets to align prosody-enriched embeddings with visual representations; (2) a 280M-parameter distilled U-Net with multi-scale fusion; and (3) a decoupled cross-attention adapter injecting prosodic embeddings via IP-Adapter. On RAVDESS (8 emotions, chance 12.5%), NovaDiffusion achieves 71.3% ECA, outperforming SonicDiffusion by 23.1pp; on held-out IEMOCAP speakers it reaches 63.4% vs. 41.7%. ECA relies on facial expression classification and does not capture scene-level emotion. Scope is limited to emotion-correlated prosody; full prosodic control remains open.

Subject: INTERSPEECH.2026 - Language and Multimodal