English

Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse Automation

Robotics 2025-07-02 v1 Computer Vision and Pattern Recognition Machine Learning

Abstract

Modern warehouse automation systems rely on fleets of intelligent robots that generate vast amounts of data -- most of which remains unannotated. This paper develops a self-supervised domain adaptation pipeline that leverages real-world, unlabeled data to improve perception models without requiring manual annotations. Our work focuses specifically on estimating the pose and shape of boxes and presents a correct-and-certify pipeline for self-supervised box pose and shape estimation. We extensively evaluate our approach across a range of simulated and real industrial settings, including adaptation to a large-scale real-world dataset of 50,000 images. The self-supervised model significantly outperforms models trained solely in simulation and shows substantial improvements over a zero-shot 3D bounding box estimation baseline.

Keywords

Cite

@article{arxiv.2507.00984,
  title  = {Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse Automation},
  author = {Xihang Yu and Rajat Talak and Jingnan Shi and Ulrich Viereck and Igor Gilitschenski and Luca Carlone},
  journal= {arXiv preprint arXiv:2507.00984},
  year   = {2025}
}

Comments

12 pages, 6 figures. This work will be presented at the 19th International Symposium on Experimental Robotics (ISER2025)

R2 v1 2026-07-01T03:42:00.541Z