English

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

Robotics 2025-01-07 v2

Abstract

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real splits, collected via simulations in MuJoCo and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so that it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/.

Keywords

Cite

@article{arxiv.2501.00510,
  title  = {VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception},
  author = {Zhaoliang Wan and Yonggen Ling and Senlin Yi and Lu Qi and Wangwei Lee and Minglei Lu and Sicheng Yang and Xiao Teng and Peng Lu and Xu Yang and Ming-Hsuan Yang and Hui Cheng},
  journal= {arXiv preprint arXiv:2501.00510},
  year   = {2025}
}
R2 v1 2026-06-28T20:53:27.463Z