English

DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation Models

Robotics 2024-02-21 v1 Machine Learning

Abstract

We propose DINOBot, a novel imitation learning framework for robot manipulation, which leverages the image-level and pixel-level capabilities of features extracted from Vision Transformers trained with DINO. When interacting with a novel object, DINOBot first uses these features to retrieve the most visually similar object experienced during human demonstrations, and then uses this object to align its end-effector with the novel object to enable effective interaction. Through a series of real-world experiments on everyday tasks, we show that exploiting both the image-level and pixel-level properties of vision foundation models enables unprecedented learning efficiency and generalisation. Videos and code are available at https://www.robot-learning.uk/dinobot.

Keywords

Cite

@article{arxiv.2402.13181,
  title  = {DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation Models},
  author = {Norman Di Palo and Edward Johns},
  journal= {arXiv preprint arXiv:2402.13181},
  year   = {2024}
}

Comments

To appear at 2024 IEEE International Conference on Robotics and Automation (ICRA)

R2 v1 2026-06-28T14:54:47.237Z