Total: 1
Recent studies indicate that the accuracy of Voice Activity Detection (VAD), a crucial component in speech processing systems, can be enhanced by leveraging both audio and visual cues. However, effectively fusing multimodal information across diverse content to achieve a highly accurate and robust VAD is still a challenge. In this paper, we present an audio-visual Modality-Aligned dual Cross-attention framework for VAD, MAC-VAD. The proposed model uses an audio encoder optimised to adaptively acquire spectral and temporal representations from raw audio. Whilst the visual encoder is designed to predict speech onsets, utilising face and lip features, accurately. The proposed model also employs knowledge distillation from a teacher network to steadily guide the multimodal framework in rigorous learning for the VAD task. Evaluations with the MMVAD dataset prove MAC-VAD's superior performance, validating the effectiveness of modality-aligned cross-attention and self-supervised distillation