tran26c@interspeech_2026@ISCA

Total: 1

#1 From Single to Multi-Label SER: Dataset and Mamba-Based Fusion Model [PDF] [Copy] [Kimi] [REL]

Authors: Vu Long Tran, Long Duc Do, Vinh Quang Nguyen, Ngoc Minh Nguyen, Quang Minh Le Pham, Hieu Trung Nguyen, Trang Thu Thi Nguyen

We present a reproducible multi-label benchmark for speech emotion recognition (SER) built from the MSP-Podcast V2.0 corpus. Our pipeline converts crowdsourced emotion votes into multi-hot labels with controlled ambiguity and sparsity via reliability filtering, deterministic vote-to-label construction, and lightweight time-shift augmentation. On this benchmark, we provide a compact Mamba-based fusion baseline over MFCC and log-mel features, following prior SER fusion practice, while using linear-time state-space backbones for efficient modeling of long audio. Using a validation-only global threshold sweep for checkpoint selection and a fixed test-time threshold for evaluation, our fusion model attains competitive micro-F1 on two MSP-Podcast test partitions (≈0.50 on both Test1 and Test2) while remaining substantially lighter than heavier baselines and recent methods. We release dataset splits, construction code, and training recipes to support future work on scalable multi-label SER.

Subject: INTERSPEECH.2026 - Speech Recognition