Total: 1
Speech Emotion Recognition (SER) benefits from large Self-Supervised Learning (SSL) models (e.g., HuBERT, WavLM), but their immense computational overhead hinders real-time edge deployment. While mel-spectrograms offer a lightweight alternative, traditional CNNs struggle to match SSL representation quality. To bridge the gap between the two approaches, we propose MSMC (Multi-Scale Masked Convolution), a highly efficient architecture for spectrogram-based SER. A masked convolution encoder (MCE) extracts sparse spectral features without information leakage, while a mean teacher framework enforces multi-scale consistency, to effectively distil global semantic context and micro-prosodic details into a compact model. Evaluated on the IEMOCAP dataset, MSMC achieves state-of-the-art accuracy among lightweight models (76.0% WA), matching the performance of heavy SSL baselines. With a significantly lower parameter count and computational complexity, MSMC is suitable for real-time SER applications.