Total: 1
We present a reproducible multi-label benchmark for speech emotion recognition (SER) built from the MSP-Podcast V2.0 corpus. Our pipeline converts crowdsourced emotion votes into multi-hot labels with controlled ambiguity and sparsity via reliability filtering, deterministic vote-to-label construction, and lightweight time-shift augmentation. On this benchmark, we provide a compact Mamba-based fusion baseline over MFCC and log-mel features, following prior SER fusion practice, while using linear-time state-space backbones for efficient modeling of long audio. Using a validation-only global threshold sweep for checkpoint selection and a fixed test-time threshold for evaluation, our fusion model attains competitive micro-F1 on two MSP-Podcast test partitions (≈0.50 on both Test1 and Test2) while remaining substantially lighter than heavier baselines and recent methods. We release dataset splits, construction code, and training recipes to support future work on scalable multi-label SER.