yanRDgQT6H@OpenReview

Total: 1

#1 TowerVision: Understanding and Improving Multilinguality in Vision-Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Andre G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi, Emmanouil Zaranis, Nuno M Guerreiro, Amin Farajian, Graham Neubig, Andre Martins

Despite rapid progress in vision-language models (VLMs), most existing approaches remain English-centric, often relying on undisclosed training data or recipes, which limits their effectiveness and reproducibility in multilingual settings. In this work, we present a systematic empirical study of how to best incorporate multilinguality across training data, encoder choices, and language models. Our results show that high-quality multilingual vision-language data substantially improve cross-lingual generalization, enabling effective transfer both from high-resource to under-represented languages and in the opposite direction. We further find that language models with strong multilingual priors are often more effective than initializing from general-purpose language models. Guided by these findings, we design TowerVision, a family of open-source multilingual VLMs, built on the multilingual text-only model Tower+. TowerVision-9B achieves competitive performance across a range of multimodal multilingual benchmarks, with particular strength in culturally grounded tasks and multi- modal translation. Notably, our models outperform existing approaches trained on substantially larger datasets, as demonstrated on ALM-Bench and Multi30K. Alongside the models, we release VISIONBLOCKS, a high-quality, curated vision-language dataset.

Subject: COLM.2026