Total: 1
Unlike Speech Emotion Recognition (SER), which predicts utterance-level labels, Speech Emotion Diarization (SED) models fine-grained temporal dynamics. However, learning reliable emotion boundaries from utterance-level labels remains challenging due to scarce frame-level annotations. We propose P-SED, a weakly supervised SED framework based on prototype metric learning that constructs a metric space via learnable emotion prototypes and introduces orthogonal regularization to enhance inter-class separability. To mitigate sparse supervision and neutral background interference, we design a Class-aware Prototype Contrastive Loss (CPCL) combined with a Top-K ranking mechanism to mine salient emotion segments. Finally, Total Variation Denoising (TVD) is applied during inference to preserve steep emotion boundaries. Experiments on the ZED dataset show that P-SED significantly outperforms existing weakly supervised models, achieving an Emotion Diarization Error Rate (EDER) of 47.00%.