2607.15794

Total: 1

#1 On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation [PDF] [Copy] [Kimi] [REL]

Authors: Stefano Silvestrini, Michele Ceresoli

Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion. This work investigates the geometric structure that emerges inside a multi-modal network for egomotion estimation. Event tensors, inertial measurements, and range signals are fused through a cross-modal attention architecture and trained in a batch setting. We analyze the latent space geometry and attention dynamics, showing that (i) embeddings lie on low-dimensional manifolds aligned with motion variables, (ii) attention weights adapt with angular excitation and visual reliability, and (iii) the fused representation recovers classical observability cues. These results bridge analytical estimation theory and modern data-driven fusion.

Subjects: Computer Vision and Pattern Recognition , Artificial Intelligence

Publish: 2026-07-17 09:58:17 UTC