Total: 1
This paper describes the FBK submissions to the Subtitling track of the 2026 IWSLT Evaluation Campaign. The task requires automatically subtitling English audio-visual content across three domains (ITV entertainment series, Asharq-Bloomberg news programs, and YouTube recordings from the YODAS dataset), into up to four target languages per domain, chosen from a pool of five (Arabic, Chinese, German, Japanese, and Spanish). All submitted systems are based on an ASR-MT cascade framework built exclusively from freely available open-source components usable without restrictions, including for commercial purposes. Our primary system implements a two-stage pipeline: the first stage produces time-aligned subtitles via voice activity detection, automatic transcription, and subtitle-level translation, while the second refinement stage re-processes the audio at a longer context level, combining long-form transcription with sentence-level translation, and re-aligning the resulting output to the original subtitle timing. This design preserves synchronization constraints while leveraging broader context to improve both transcription and translation quality. We also submitted two contrastive systems: one corresponding to the first-stage baseline pipeline, and another sharing the same baseline architecture but using alternative components.