Total: 1
Aligning non-invasive neural activity with speech stimuli is a foundational challenge in neural speech decoding. The difficulty lies in mapping noisy, high-dimensional electroencephalography (EEG) signals to the temporal dynamics of speech features. We propose SCANS, a supervised contrastive learning framework for neural-speech temporal alignment. We define this as a classification task where, given an EEG segment and multiple non-overlapping candidate segments from the same speech signal, the model must identify the single matching stimulus temporally aligned with the EEG. SCANS utilizes a dilated convolutional frontend and cross-modal attention to extract and fuse features across modalities. To bridge the modality gap, we employ a multi-task objective combining cross-entropy classification with a contrastive loss. Evaluated on the SParrKULee dataset, SCANS demonstrates significant improvement in neural-speech alignment accuracy.