2509.20579

Total: 1

#1 Large Pre-Trained Models for Bimanual Manipulation in 3D [PDF3] [Copy] [Kimi1] [REL]

Authors: Hanna Yurchyk, Wei-Di Chang, Gregory Dudek, David Meger

We investigate the integration of attention maps from a pre-trained Vision Transformer into voxel representations to enhance bimanual robotic manipulation. Specifically, we extract attention maps from DINOv2, a self-supervised ViT model, and interpret them as pixel-level saliency scores over RGB images. These maps are lifted into a 3D voxel grid, resulting in voxel-level semantic cues that are incorporated into a behavior cloning policy. When integrated into a state-of-the-art voxel-based policy, our attention-guided featurization yields an average absolute improvement of 8.2% and a relative gain of 21.9% across all tasks in the RLBench bimanual benchmark.

Subjects: Computer Vision and Pattern Recognition , Machine Learning , Robotics

Publish: 2025-09-24 21:38:42 UTC