English

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Computer Vision and Pattern Recognition 2025-09-09 v1 Artificial Intelligence Robotics

Abstract

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.

Keywords

Cite

@article{arxiv.2509.05513,
  title  = {OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation},
  author = {Ahad Jawaid and Yu Xiang},
  journal= {arXiv preprint arXiv:2509.05513},
  year   = {2025}
}

Comments

4 pages, 1 figure

R2 v1 2026-07-01T05:23:57.302Z