Total: 1
Speech emotion recognition (SER) faces challenges due to label ambiguity and high model uncertainty during early training. Standard training with hard one-hot targets ignores these issues, forcing models to commit to a single class before learning meaningful representations. We propose Progressive Weak Supervision (PWS), a strategy that relaxes supervision in the early stages and gradually tightens it as the model matures. During the initial epochs, a sample is considered correctly predicted if the true label is among the top-k predictions, with a soft probability distribution assigned across them. As training progresses, k decays linearly to 1, thereby recovering standard supervision. We evaluate PWS using a WavLM-Base encoder on IEMOCAP (English, 4-class) and ViSEC (Vietnamese, 4-class). With k=3, PWS achieves 78.08% unweighted accuracy on IEMOCAP and 85.70% on ViSEC, surpassing several strong baseline methods. Code is available at https://github.com/skyemo47/PWS.