2608.16360

Total: 1

#1 Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding [PDF1] [Copy] [Kimi] [REL]

Authors: Tomoaki Mizuno, Toru Nakashika

Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.

Subjects: Audio and Speech Processing , Sound , Signal Processing

Publish: 2026-08-17 10:11:11 UTC