OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

#1 OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment [PDF¹] [Copy] [Kimi] [REL]

Authors: Tianchao Li, Shujian Yu, Xinrui Zu, Zhaolong Wei, Jeremy Gummeson, Jack C. P. Cheng, Robert Jenssen

Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent two-way comparisons and therefore do not explicitly model higher-order dependencies among multiple modalities. Recent beyond-pairwise objectives approach this problem from statistical or geometric perspectives, but arbitrary-modality alignment still lacks a principled criterion for defining what each modality should preserve and compress relative to the others. We revisit arbitrary-modality alignment through the Information Bottleneck principle. In multi-modal learning, sufficiency should preserve information predictable from the remaining modalities, while minimality should compress modality-specific information not supported by them. This naturally leads to a One-vs-All view, where each modality is characterized with respect to the remaining modalities. We propose OVA-IB, an Information Bottleneck framework for arbitrary-modality alignment. OVA-IB optimizes a tractable One-vs-All contrastive lower bound for sufficiency connected to a Dual Total Correlation-style objective, uses a parameter-free geometry-aware projection score, and derives a tractable upper-bound regularizer for minimality by bounding each representation's dependence on its own input with representation distributions induced by the remaining modalities. Experiments on classification, regression, modality-agnostic evaluation, and cross-modal retrieval benchmarks demonstrate strong and robust performance.

Subjects: Machine Learning , Information Theory

Publish: 2026-05-28 13:23:07 UTC

2605.29900

#1 OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment [PDF1] [Copy] [Kimi] [REL]

#1 OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment [PDF¹] [Copy] [Kimi] [REL]