Total: 1
A Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework was previously proposed to integrate speaker diarization with speech separation for handling the realistic multi-speaker scenarios. However, DCF-DS mainly relies on spectral cues to distinguish speakers, and its performance degrades in highly overlapped regions. In this paper, we first extend DCF-DS to multi-channel DCF-DS (MC-DCF-DS) by incorporating spatial information at the system level. We then further propose MCA-DCF-DS by introducing spatial information into the adaptation data simulation process to handle the trade-off between miss and confusion errors in diarization results. Experimental results demonstrate that incorporating spatial information at both the system and data levels improves downstream ASR performance. Moreover, under the same ASR backend, the proposed MCA-DCF-DS outperforms the CHiME-8 Task 2 champion system.