lan26@interspeech_2026@ISCA

Total: 1

#1 SA-UAED: Joint Frame-Level Detection of Audio Events, Speaker Activities, and Speaker-Attributed Paralinguistic Events [PDF] [Copy] [Kimi] [REL]

Authors: Zekun Lan, Wangyou Zhang, Yanmin Qian

Existing sound event detection (SED) and speaker diarization (SD) systems typically fail to attribute paralinguistic events (e.g., laughter, coughing) to individual speakers, primarily due to the scarcity of fine-grained labeled data. To address this, we propose a dedicated simulation pipeline generating large-scale audio mixtures with precise frame-level annotations. Leveraging this, we introduce the Speaker-Attributed Unified Audio Event Detection (SA-UAED) framework for joint frame-level prediction of sound events, speaker activities, and paralinguistics. Experiments demonstrate that SA-UAED substantially improves speaker-attributed paralinguistic detection without compromising generic SED and SD accuracy.

Subject: INTERSPEECH.2026 - Modelling and Learning