Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

duquenne23@interspeech_2023@ISCA

Total: 1

#1 Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer [PDF²] [Copy] [Kimi²] [REL]

Authors: Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

Recent research has shown that independently trained encoders and decoders, combined through a shared fixed-size representation, can achieve competitive performance in speech-to-text translation. In this work, we show that this type of approach can be further improved with multilingual training. We observe significant improvements in zero-shot cross-modal speech translation, even outperforming a supervised approach based on XLSR for several languages.

Subject: INTERSPEECH.2023 - Language and Multimodal

duquenne23@interspeech_2023@ISCA

#1 Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer [PDF2] [Copy] [Kimi2] [REL]

#1 Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer [PDF²] [Copy] [Kimi²] [REL]