IWSLT.2026

| Total: 39

#1 Towards Zero-Shot SLU: An Empirical Study of Competing Architectural Paradigms [PDF1] [Copy] [Kimi1] [REL]

Authors: Beomseok Lee, Marco Gaido, Ioan Calapodescu, Laurent Besacier, Matteo Negri

Spoken Language Understanding (SLU) is crucial for enabling natural voice interactions with modern devices. However, traditional supervised models fail to generalize to new domains due to two key challenges: the prohibitive cost of data annotation and the inherent difficulty of transferring domain-specific intents. While the rise of Large Language Models (LLMs) offers a promising solution through zero-shot inference, the zero-shot SLU capabilities of emerging speech-enabled LLMs have remained largely unexplored. To address this gap, this paper provides the first comprehensive assessment, focusing on intent classification (IC), the first key sub-task of SLU, across 13 languages. We systematically evaluate a range of architectures, including cascaded, end-to-end, and hybrid systems for zero-shot SLU. Our analysis identifies the hybrid approach as the most effective architectural design for end-to-end SLU, and assesses multilingual transfer capabilities. The findings offer a detailed map of the challenges and opportunities, highlighting which models and settings are most promising for zero-shot SLU.

Subject: IWSLT.2026


#2 Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies [PDF] [Copy] [Kimi] [REL]

Authors: Qianen Zhang, Zeyu Yang, Satoshi Nakamura

Simultaneous Machine Translation (SiMT) requires high-quality translations under strict real-time constraints, which traditional policies with only READ/WRITE actions cannot fully address. We extend the action space of SiMT with four adaptive actions: Sentence_Cut, Drop, Partial_Summarization and Pronominalization, which enable real-time restructuring, omission, and simplification while preserving semantic fidelity. We adapt these actions in a large language model (LLM) framework and construct training references through action-aware prompting. To evaluate both quality and word-level monotonicity, we further develop a latency-aware TTS pipeline that maps textual outputs to speech with realistic timing. Experiments on the ACL60/60 English-Chinese, English-German and English-Japanese benchmarks show that our framework consistently improves semantic metrics and achieves lower delay compared to reference translations and salami-based baselines. Notably, combining Drop and Sentence_Cut leads to consistent improvements in the balance between fluency and latency. These results demonstrate that enriching the action space of LLM-based SiMT provides a promising direction for bridging the gap between human and machine interpretation.

Subject: IWSLT.2026


#3 A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation [PDF] [Copy] [Kimi] [REL]

Authors: Yulin Xue, Siqi Ouyang, Lei Li

Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. Prior approaches are difficult to reproduce and make assumptions that do not hold for end-to-end systems. We present a practical evaluation method for long-form SimulS2ST. Given source speech, pre-segmented source transcripts, and reference translations, we run automatic speech recognition (ASR) and forced alignment on the generated target speech to recover token-level timestamps, then apply a sentence-embedding-based aligner to match the target text to its corresponding source sentences. This enables sentence-level computation of latency and quality metrics, including YAAL and xCOMET, which are then aggregated into final system-level scores. Experiments on representative SimulS2ST systems show that the method is effective in practice and reveal that current systems suffer from substantial latency accumulation on long speech. Code can be found here https://github.com/SakaiXue6666/Speech-to-Speech-Latency

Subject: IWSLT.2026


#4 Selected-Layer Codec Compression for Compact Speech Translation Models: An IWSLT 2026 English-to-Chinese Submission [PDF] [Copy] [Kimi] [REL]

Author: Alonso Palomino

This paper describes a selected-layer codec compression approach submitted to the IWSLT 2026 Model Compression Shared Task for constrained English-to-Chinese speech translation. The approach is compared against standard quantization, global codec compression, and a pruning-plus-codec variant. The results indicate that translation quality after compression depends strongly on where compression is applied. In these experiments, selected-layer compression preserves translation quality better than uniform global compression, with one variant achieving the highest COMET score among compressed systems and another providing the strongest overall quality-compression trade-off among the custom codec methods. These results suggest that simple layer-aware post-hoc compression is a viable approach for model compression in constrained English-to-Chinese speech translation.

Subject: IWSLT.2026


#5 Team QUESPA System Submission for the IWSLT 2026 Dialectal and Low-resource Speech Translation Task [PDF] [Copy] [Kimi] [REL]

Authors: John E. Ortega, Rodolfo Joel Zevallos, Fabrício Carraro, Stephanny Gabriela Sánchez Bautista, Chad Howe

This paper describes the QUESPA team’s speech translation (ST) submissions for the Quechua to Spanish (QUE-SPA) track of the IWSLT 2026 Evaluation Campaign on dialectal and low-resource speech translation. The campaign supports a single submission category, namely unconstrained. This marks our fourth consecutive participation in the IWSLT shared task, building upon prior systems with substantial improvements. Our 2026 submission comprises three unconstrained-only systems. The best-performing system (contrastive 2) extends our strongest model from the previous year by leveraging a high-performing pre-trained language model (PLM) for end-to-end speech translation without cascading, augmented with additional Quechua-Collao text - now made available on the IWSLT GitHub. Fine-tuning Microsoft’s SpeechT5 model in an ST setting, combined with targeted data augmentation, results in a BLEU score of 27.2 on the official evaluation set. Additionally, we evaluate prompt-based machine translation using Gemini, DeepSeek, GPT-5, Claude, and Qwen for the first time. Aside from that, we introduce SIDON, an audio enhancement framework designed to improve audio quality. This paper provides a comparative analysis across our current and three previous IWSLT submissions, with a detailed examination of the impact of synthetic data, unconstrained external resources, and audio enhancement techniques on fine-tuning performance. Our results highlight the complementary role of PLM-based ST, LLM prompting, and ASR enhancement in advancing low-resource speech translation.

Subject: IWSLT.2026


#6 ADAPT–MTU HAI at IWSLT2026: Robust Cascaded Speech Translation for Bhojpuri–Hindi and Irish–English [PDF] [Copy] [Kimi] [REL]

Authors: Pournima Sonawane, Haithem Afli

Low-resource speech translation remains challenging due to limited data, weak ASR support, and error propagation in cascaded systems. We present the ADAPT–MTU HAI submission to the IWSLT 2026 Low-Resource Speech Translation task, a robust cascaded framework combining Whisper-based ASR and NLLB-200 multilingual translation for Bhojpuri→Hindi and Irish→English language pairs. We evaluate multiple ASR models and routing strategies, including direct and pivot-based translation. For Bhojpuri→Hindi, the best configuration (Whisper-large-v3 and direct NLLB) achieves BLEU 25.59, chrF++ 42.48, and TER 63.83 on the full development set, outperforming pivot and copy baselines. For Irish→English, replacing Whisper with a language-specific Wav2Vec2 ASR model improves ASR coverage from 94.8% to 100% on the test set while maintaining low repetition rates. Our findings highlight the critical role of ASR quality in downstream translation performance, the conditional benefits of pivot translation, and the effectiveness of modular cascaded architectures for low-resource speech translation.

Subject: IWSLT.2026


#7 The FBK Sentence-Aware Subtitling System at the IWSLT 2026 Subtitling Track [PDF] [Copy] [Kimi] [REL]

Authors: Mauro Cettolo, Roldano Cattoni, Matteo Negri, Luisa Bentivogli

This paper describes the FBK submissions to the Subtitling track of the 2026 IWSLT Evaluation Campaign. The task requires automatically subtitling English audio-visual content across three domains (ITV entertainment series, Asharq-Bloomberg news programs, and YouTube recordings from the YODAS dataset), into up to four target languages per domain, chosen from a pool of five (Arabic, Chinese, German, Japanese, and Spanish). All submitted systems are based on an ASR-MT cascade framework built exclusively from freely available open-source components usable without restrictions, including for commercial purposes. Our primary system implements a two-stage pipeline: the first stage produces time-aligned subtitles via voice activity detection, automatic transcription, and subtitle-level translation, while the second refinement stage re-processes the audio at a longer context level, combining long-form transcription with sentence-level translation, and re-aligning the resulting output to the original subtitle timing. This design preserves synchronization constraints while leveraging broader context to improve both transcription and translation quality. We also submitted two contrastive systems: one corresponding to the first-stage baseline pipeline, and another sharing the same baseline architecture but using alternative components.

Subject: IWSLT.2026


#8 KIT’s Submission to Cross-Lingual Voice Cloning in IWSLT 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Seymanur Akti, Alexander Waibel

Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translation and is the focus of the IWSLT 2026 Cross-Lingual Voice Cloning track. A key challenge is maintaining intelligibility and naturalness in the presence of accent variation and domain-specific vocabulary. We build on a multilingual text-to-speech model, FishAudio-S2-Pro, and introduce language tag prompting to improve language control and reduce accent leakage. We further apply reinforcement learning (RL) fine-tuning for task adaptation and observe improvements in intelligibility. Finally, we propose a reference-conditioned lexical matching method that improves pronunciation of domain-specific terms when lexical overlap is present. Results show that language prompting provides the largest gains, while lexical matching yields consistent improvements on matched subsets.

Subject: IWSLT.2026


#9 HW-TSC’s Submissions to the IWSLT 2026 Offline Speech Translation Task [PDF] [Copy] [Kimi] [REL]

Authors: Boqi Huang, Daimeng Wei, Jiaxin GUO, Yuanchang Luo, Hengchao Shang, Zongyao Li, Zhiqiang Rao, Jinlong Yang, Zhanglin Wu, Yu He, Xiaoqing Lan

This paper describes the HW-TSC’s submission to the IWSLT 2026 Offline Speech Translation Task, specifically for the English-to-Chinese and English-to-German unconstrained tracks. Our system adopts a robust cascade architecture optimized for long-form, unsegmented audio. To mitigate the hallucination and inconsistency issues common in long-sequence processing, we propose a two-pass transcription strategy: an initial streaming ASR with a 12-second context buffer for sentence-level coherence, followed by Qwen3-ForcedAligner for precise timestamping. Based on these alignments, a second-pass refinement is conducted using Qwen3-Omni on re-segmented 30-second chunks to ensure high-fidelity transcriptions. For the translation module, we employ a context-aware segment merging strategy (up to 150 tokens) to empower the Qwen3 llm with sufficient semantic context. Experimental results on the tst-2022 benchmark demonstrate the effectiveness of our pipeline, achieving COMET scores of 0.8462 (En-Zh) and 0.7854 (En-De), significantly outperforming the standard cascade baselines.

Subject: IWSLT.2026


#10 HW-TSC’s Submission to the IWSLT 2026 Subtitling Track [PDF] [Copy] [Kimi] [REL]

Authors: Xiaoqing Lan, Daimeng Wei, Jiaxin GUO, Yuanchang Luo, Hengchao Shang, Zongyao Li, Zhiqiang Rao, Jinlong Yang, Zhanglin Wu, Boqi Huang, Yu He

This paper introduces HW-TSC’s submission to the IWSLT 2026 Subtitling track. For automatic subtitle generation, we employ a cascaded strategy under unconstrained conditions. First, we construct a large-model-based streaming speech recognition framework, which incorporates VAD voice activity detection, sliding-window context caching, long audio chunking, and the Qwen3 forced alignment model to achieve high-precision transcription and timestamping from English speech to text. Next, we perform text translation using a Qwen3-based translation model. Finally, according to subtitle constraints such as characters per second (CPS) and characters per line (CPL), we identify translation segments that exceed compliance thresholds via quantitative evaluation, and rewrite them using a large language model while preserving core semantic meaning, ultimately producing subtitle files that meet the required standards.

Subject: IWSLT.2026


#11 HW-TSC’s Submission to the IWSLT 2026 Cross-Lingual Voice Cloning Track [PDF] [Copy] [Kimi] [REL]

Authors: Yu He, Daimeng Wei, Jiaxin GUO, Yuanchang Luo, Hengchao Shang, Zongyao Li, Zhiqiang Rao, Jinlong Yang, Zhanglin Wu, Boqi Huang, Xiaoqing Lan

This paper presents HW-TSC’s submission to the IWSLT 2026 Cross-Lingual Voice Cloning Track. The Cross-Lingual Voice Cloning Track includes three target languages: Arabic, Chinese, and French. We take part in two language tasks of this track, namely Chinese and French. We employ the Qwen3-TTS-12Hz-1.7B-Base multilingual model as the core voice cloning model. To tackle problems such as excessively long duration of the original reference audio and scattered features, we design a sliding-window audio segmentation preprocessing method, which continuously splits long audio into standardized short segments with overlapping redundancy. This method avoids feature attenuation caused by overly long audio and maximizes the preservation of complete timbre information through step overlap. To select the outputs with the highest timbre similarity from numerous synthetic results, this study conducts voiceprint recognition based on the Enhanced Context-Dependent Adversarial Time Delay Neural Network (ECAPA-TDNN), with cosine similarity as the core quantitative evaluation metric, and selects the result with the highest similarity as the optimal output.

Subject: IWSLT.2026


#12 Balancing Linguistic Intelligibility and Speaker Identity in Zero-Shot Cross-Lingual Voice Cloning [PDF] [Copy] [Kimi] [REL]

Authors: Mo Ahtasam, Jamal uddin, Mohammad Nadeem

Cross-lingual voice cloning (CLVC) aims to synthesize speech in a target language while preserving the vocal identity of a source speaker who has no recorded speech in that language. Despite recent advances in multilingual text-to-speech systems, zero-shot CLVC remains challenging due to phonetic divergence across languages and the difficulty of maintaining speaker identity alongside linguistic intelligibility. In this work, we present a systematic evaluation of four state-of-the-art CLVC systems spanning autoregressive and diffusion-based architectures. Using English source speakers from the ACL-60/60 dataset, we evaluate zero-shot voice transfer across multiple target languages, including Arabic, Chinese, French, German, Russian, and Japanese. Systems are assessed using speaker similarity and content consistency metrics under a unified multilingual evaluation pipeline. We analyze how different modeling approaches autoregressive language modeling and diffusion-based flow matching handle the tradeoff between speech accuracy and speaker identity preservation across different architectural approaches. We further observe substantial performance variation across languages, with Arabic remaining particularly challenging under zero-shot transfer settings.

Subject: IWSLT.2026


#13 CUHKSZ Simultaneous Speech Translation System for IWSLT 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Zeyu Yang, Satoshi Nakamura

We present the CUHKSZ Team submission to the IWSLT 2026 Simultaneous Speech Translation evaluation, targeting the main and Extra Context tracks for English→Chinese, German on unsegmented speech. Our system is built upon Qwen3-Omni-30B-A3B, a natively aligned audio-text LLM. Under the Constrained condition, we apply LoRA adaptation exclusively to the LLM. Specifically, we construct syntax-aware, chunk-aligned supervision from existing ASR corpora, using Qwen3-30B-Instruct to synthesize target translations. This enables the model to internalize the simultaneous read/write policy by autonomously predicting <wait> tokens at semantically incomplete boundaries. With the policy internalized, execution is delegated to a lightweight streaming agent served via vLLM. This agent feeds audio in fixed chunks, manages a bounded dialogue history, and enforces strict emission controls to minimize computation-aware delay. For the sub-track, contextual priors are dynamically injected into the prompt. On the official dev set, our 0–2 s latency regime submissions achieve 40.5 BLEU (1.95 s) for En→Zh and 27.7 BLEU (1.72 s) for En→De. In the 2–4 s regime, performance scales to 42.1 BLEU (2.16 s) and 30.5 BLEU (2.29 s) respectively.

Subject: IWSLT.2026


#14 Fleurs-Badini: Translation and Recording Fleurs Dataset for Badini Variant of Northern Kurdish [PDF] [Copy] [Kimi] [REL]

Authors: Mohammad Mohammadamini, Dilgash Mohammed Salih Tayib, Dezheen H. Abdulazeez, Barzan Hussein Mohammed, Imad Saeed Sadeeq, Aveen Jalal Mohammed, Amera Ismail Melhum, Abuobaida Abdullah Dheyab

Multilingual speech benchmarks such as the FLEURS benchmark have significantly advanced research across a wide range of languages. However, important dialects, including Badini Kurdish, remain underrepresented, limiting bechmarking in automatic speech recognition (ASR) and speech-to-text translation (S2TT). To address this limitation, this study introduces FLEURS-Badini, a dialect-focused extension designed to support research on Northern Kurdish (Badini). The dataset is constructed through a structured process of translation, recording, and validation, resulting in 5,224 utterances paired with their corresponding translated text. The data were collected from 45 speakers. To evaluate the dataset, baseline experiments are conducted using state-of-the-art models for both ASR and S2TT. The results indicate that ASR remains challenging, with the best performance achieved by the W2V-BERT CTC model, reaching a Word Error Rate (WER) of approximately 55% on the test set. Similarly, speech-to-text translation performance is limited, with BLEU scores 6.13 and 5.24 on dev and test sets. Overall, FLEURS-Badini expands multilingual coverage and provides a standardized foundation for evaluating ASR and speech translation systems in the Badini dialect.

Subject: IWSLT.2026


#15 LIUM Submission for IWSLT 2026 Low-resource Speech Translation Track [PDF] [Copy] [Kimi] [REL]

Authors: Mohammad Mohammadamini, Marie Tahon

This paper describes the LIUM submission to the IWSLT 2026 low-resource speech translation track. It proposes different data augmentation methods for low-resource speech-to-text translation, including two main pipelines: pseudo-labeling and speech synthesis. The goal is to generate parallel speech data in low-resource scenarios without relying on human-annotated speech translation data. Our submission focuses on Central Kurdish–English language pairs. The objective of this work is to explore the advantages and limitations of each data augmentation method. Our best results are obtained using the pseudo-labeling pipeline, achieving a BLEU score of 25.73 on the development set and 21.09 on the test set for Central Kurdish–English translation.

Subject: IWSLT.2026


#16 Multilingual Long-Form Speech Instruction Following: KIT’s Submission to IWSLT 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Enes Yavuz Ugan, Maike Züfle, Yuka Ko, Supriti Sinhamahapatra, Fabian Retkowski, Seymanur Akti, Jan Niehues, Alexander Waibel

With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly from natural language prompts. This trend is reflected in IWSLT’s Instruction Following Track, which this year introduced new tasks including an unknown surprise task, posing a genuine challenge against overfitting to known tasks. We present KIT’s submission to the Long and Short Instruction Following tracks in the unconstrained setting. Our approach combines a general data augmentation pipeline that converts short-form corpora into long-form training data through segment concatenation, LLM-based label generation, and cross-lingual translation, yielding over 1M instances across six tasks and four languages. We further show that likelihood-based re-ranking, while highly effective for ASR, systematically degrades semantic tasks by spuriously selecting candidates generated from segmented audio processing rather than holistic long-form inference, a failure mode resolved by combining likelihood with Minimum Bayes Risk decoding.

Subject: IWSLT.2026


#17 NAVER LABS Europe Submission to the Instruction-following 2026 Short Track [PDF] [Copy] [Kimi] [REL]

Authors: Marcely Zanon Boito, Hemant Yadav, Jean-Luc Meunier, Ioan Calapodescu

In this paper, we describe NAVER LABS Europe’s submission to the instruction-following speech processing short track at IWSLT 2026. We participate again in the constrained setting, developing systems capable of jointly performing ASR, ST, and SQA from English speech into Chinese, Italian, and German. Building on our previous submission, ranked first in last year’s short track, we update our multi-stage training pipeline by replacing the speech projector with SpeechMapper, a method for learning a speech-to-LLM embedding projector using ASR-only data. In addition, we introduce a synthetic SQA dataset, fakACL, composed of artificially generated scientific presentations. This dataset is built by prompting the LLM backbone, segmenting the generated talks, and synthesizing speech with Seamless. The combination of an improved speech projection mechanism and domain-specific synthetic data allows our model to outperform last year’s best short-track system, while being considerably more compact and relying on a weaker LLM backbone.

Subject: IWSLT.2026


#18 CATENG Submission for the IWSLT 2026: Dialectal and Low-resource Speech Translation Task [PDF] [Copy] [Kimi] [REL]

Authors: Rodolfo Joel Zevallos, Marc Casals, John E. Ortega, Fabrício Carraro, Pol Buitrago, Guillermo Cámbara

We present the CATENG systems submitted to the IWSLT 2026 Dialectal and Low-Resource Speech Translation shared task for the Catalan–English (CA–EN) pair. Although Catalan is not strictly low-resource, its dialectal diversity and relative under-representation in speech technology make it a challenging setting. We evaluate three unconstrained systems: two cascaded approaches combining ASR and MT, and one end-to-end model. Our primary system uses a Mamba-based ASR (ConMamba) with a fine-tuned NLLB-200 MT model, while a contrastive system replaces the ASR with Whisper-v3; we also evaluate an end-to-end SpeechT5 model with data augmentation. Experiments are conducted on the IWSLT 2026 Catalan dataset (15 hours), complemented with large-scale parallel text. Results show that cascaded systems outperform end-to-end ST, with Whisper-v3 + NLLB achieving 44.7 BLEU and 65.1 chrF. We find that performance is primarily constrained by ASR quality rather than MT capacity, and that Mamba-based ASR models provide competitive results, highlighting the importance of robust speech representations and dialectal coverage for Catalan–English speech translation.

Subject: IWSLT.2026


#19 BSC’s Submission to the Instruction Following Track of IWSLT 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Oriol Pareras, Joan Llado, Pol Buitrago, Marc Casals-Salvador, Federico Costa, Cristina Espana-Bonet

We present the Barcelona Supercomputing Center (BSC) submission to the Instruction Following (IF) track of IWSLT 2026, which evaluates unified spoken language systems capable of solving multiple tasks through natural language instructions. Our system consists of an end-to-end (E2E) architecture that combines a speech encoder with a translation-oriented Large Language Model. The model is trained on speech and text data, covering automatic speech recognition, translation, question answering, and instruction following. We investigate a Chain-of-Thought (CoT) generation strategy that explicitly decomposes tasks by producing an intermediate transcription before the final output, which enables effective reuse of text-only supervision and improves robustness across tasks. To further support generalization, we design diverse prompt formulations and align text-only and speech inputs under a shared inference pattern. Results on IWSLT 2025 evaluation data show that our approach achieves competitive and even state-of-the-art performance across tasks.

Subject: IWSLT.2026


#20 Towards Dynamic Attention Masking for Simultaneous Speech Translation [PDF] [Copy] [Kimi] [REL]

Author: Benjamin Pong

We present a proof-of-concept system for simultaneous speech translation based on dynamic attention masking. Our approach builds on SeamlessM4T by injecting lightweight per-layer schedulers into the conformer-encoder, training each scheduler to predict the number of future frames needed for translation. The schedulers are trained jointly with LoRA adapters across three language directions: English to German, Italian, and Chinese. At inference time, we evaluate our system using sliding window retranslation inference regime (Sen et al., 2022), and an adapted version of StreamAtt (Papi et al., 2024) that replaces the fixed cutoff with a content-aware threshold derived from the learnt representations from the scheduler outputs.

Subject: IWSLT.2026


#21 Diet-KIT: Post-Training Quantization for Speech LLMs [PDF] [Copy] [Kimi] [REL]

Authors: Danni Liu, Sai Koneru, Jan Niehues

We present Diet-KIT, a system for the IWSLT speech translation compression task under a strict 4 GB on-disk storage constraint, starting from the 16 GB Qwen2-Audio-7B base model. Compression is achieved with a sequential pipeline based on Half-Quadratic Quantization (HQQ). Based on systematic ablations, we find that 4-bit quantization preserves translation quality well, whereas 3-bit quantization induces a sharp performance cliff, precluding aggressive compression across the whole model. We further show that the embedding table tolerates 2-bit quantization with negligible loss, while the LM head requires higher precision. To satisfy the storage constraint, we propose a sensitivity-guided layer selection method that identifies MLP sublayers tolerant to 3-bit compression via a per-layer sensitivity analysis, which consistently outperforms manual and random layer selection. Finally, AWQ calibration is applied as a data-driven refinement stage. The final system achieves 3.98 GB on disk with COMET scores of 74.4 on en→de and 77.1 on en→zh, compared to 75.6 and 79.5 for the uncompressed fine-tuned model.

Subject: IWSLT.2026


#22 A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Aziz Sharipov Ortega, Dominik Macháček

We implement a direct speech translation model Canary for simultaneous translation with AlignAtt simultaneous policy. We focus on Nemo toolkit with the recent state-of-the-art foundation model Canary-1B-v2 that has only one billion of parameters, which is suitable for small pocket devices. This is a CUNI submission to IWSLT 2026 Simultaneous Speech Translation Shared task on Czech to English and English to German and Italian.

Subject: IWSLT.2026


#23 NeMo@IWSLT 2026: Cascaded System for Simultaneous Speech Translation [PDF] [Copy] [Kimi] [REL]

Authors: Lilit Grigoryan, Vladimir Bataev, Andrei Andrusenko, Oleksii Hrinchuk, Davit Karamyan, Enas Albasiri, Vitaly Lavrukhin, Nikolay Karpov, Boris Ginsburg

This paper describes the NVIDIA NeMo team’s submission to the IWSLT 2026 Simultaneous Speech Translation (SimulST) tracks. We use a cascaded architecture combining a dual-mode Unified ASR Transducer model with a multilingual Large Language Model (LLM). The ASR is trained to deliver stable transcriptions across wide range of latencies, providing a reliable foundation for high-quality LLM translation. Our submission participates in the English–German, English–Italian, and English–Chinese tasks, in both standard and contextualized settings, as well as the Czech–English standard track, covering both low- and high-latency scenarios. We further analyze how ASR and LLM design choices affect the system’s overall latency and translation quality.

Subject: IWSLT.2026


#24 MLLP-VRAIN UPV System for the IWSLT 2026 Simultaneous Speech Translation Task [PDF] [Copy] [Kimi] [REL]

Authors: Jorge Iranzo-Sánchez, Gerard Mas-Mollà, Adrià Gimenez, Jorge Civera Saiz, Albert Sanchis, Alfons Juan

This work describes the participation of the MLLP-VRAIN research group in the shared task of the IWSLT 2026 Simultaneous Speech Translation track. Our submission utilizes the recently released Parakeet and Qwen 3.5 models to create a robust, cascaded solution for long-form SimulST through the use of adaptive black-box policies. We explore relaxations of these policies to achieve better quality-latency trade-offs. Compared to last year, we participate on all language directions. In addition to this, for the En→De, It, Zh directions we also participate in this year’s new context track employing a combination of ASR word-boosting and a RAG mechanism of offline pre-translated exemplars to guide generation and enrich our system with domain-specific context. Finally, we provide a detailed latency analysis of our system. Compared to last year, results on the MCIF En→De test set shows a substantial quality improvement of +5.82 XCOMET-XL. Our context track processing further improves performance by +1.03.

Subject: IWSLT.2026


#25 One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech [PDF] [Copy] [Kimi] [REL]

Authors: Amanuel Gizachew Abebe, Yasmin Moslem

Preserving a speaker’s voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we address this challenge through our system submission to the International Conference on Spoken Language Translation (IWSLT 2026), the Cross-Lingual Voice Cloning shared task. First, we evaluate several state-of-the-art voice cloning models for cross-lingual speech generation of scientific texts in Arabic, Chinese, and French. Then, we build voice cloning systems based on the OmniVoice foundation model. We employ data augmentation via multi-model ensemble distillation from the ACL 60/60 corpus. We investigate the effect of using this synthetic data for fine-tuning, demonstrating improvements in intelligibility (WER and CER) and speaker similarity (SIM), with gains varying across languages.

Subject: IWSLT.2026