ta26b@interspeech_2026@ISCA

Total: 1

#1 Progressive Weak Supervision for Speech Emotion Recognition [PDF] [Copy] [Kimi] [REL]

Authors: Bao Thang Ta, Huynh Thi Thanh Binh, Van Hai Do

Speech emotion recognition (SER) faces challenges due to label ambiguity and high model uncertainty during early training. Standard training with hard one-hot targets ignores these issues, forcing models to commit to a single class before learning meaningful representations. We propose Progressive Weak Supervision (PWS), a strategy that relaxes supervision in the early stages and gradually tightens it as the model matures. During the initial epochs, a sample is considered correctly predicted if the true label is among the top-k predictions, with a soft probability distribution assigned across them. As training progresses, k decays linearly to 1, thereby recovering standard supervision. We evaluate PWS using a WavLM-Base encoder on IEMOCAP (English, 4-class) and ViSEC (Vietnamese, 4-class). With k=3, PWS achieves 78.08% unweighted accuracy on IEMOCAP and 85.70% on ViSEC, surpassing several strong baseline methods. Code is available at https://github.com/skyemo47/PWS.

Subject: INTERSPEECH.2026 - Speech Recognition