English

Large Pre-Trained Models for Bimanual Manipulation in 3D

Computer Vision and Pattern Recognition 2025-11-13 v1 Machine Learning Robotics

Abstract

We investigate the integration of attention maps from a pre-trained Vision Transformer into voxel representations to enhance bimanual robotic manipulation. Specifically, we extract attention maps from DINOv2, a self-supervised ViT model, and interpret them as pixel-level saliency scores over RGB images. These maps are lifted into a 3D voxel grid, resulting in voxel-level semantic cues that are incorporated into a behavior cloning policy. When integrated into a state-of-the-art voxel-based policy, our attention-guided featurization yields an average absolute improvement of 8.2% and a relative gain of 21.9% across all tasks in the RLBench bimanual benchmark.

Keywords

Cite

@article{arxiv.2509.20579,
  title  = {Large Pre-Trained Models for Bimanual Manipulation in 3D},
  author = {Hanna Yurchyk and Wei-Di Chang and Gregory Dudek and David Meger},
  journal= {arXiv preprint arXiv:2509.20579},
  year   = {2025}
}

Comments

Accepted to 2025 IEEE-RAS 24th International Conference on Humanoid Robots

R2 v1 2026-07-01T05:55:00.952Z