Total: 1
Existing sound event detection (SED) and speaker diarization (SD) systems typically fail to attribute paralinguistic events (e.g., laughter, coughing) to individual speakers, primarily due to the scarcity of fine-grained labeled data. To address this, we propose a dedicated simulation pipeline generating large-scale audio mixtures with precise frame-level annotations. Leveraging this, we introduce the Speaker-Attributed Unified Audio Event Detection (SA-UAED) framework for joint frame-level prediction of sound events, speaker activities, and paralinguistics. Experiments demonstrate that SA-UAED substantially improves speaker-attributed paralinguistic detection without compromising generic SED and SD accuracy.