English

TexHOI: Reconstructing Textures of 3D Unknown Objects in Monocular Hand-Object Interaction Scenes

Computer Vision and Pattern Recognition 2025-05-12 v2

Abstract

Reconstructing 3D models of dynamic, real-world objects with high-fidelity textures from monocular frame sequences has been a challenging problem in recent years. This difficulty stems from factors such as shadows, indirect illumination, and inaccurate object-pose estimations due to occluding hand-object interactions. To address these challenges, we propose a novel approach that predicts the hand's impact on environmental visibility and indirect illumination on the object's surface albedo. Our method first learns the geometry and low-fidelity texture of the object, hand, and background through composite rendering of radiance fields. Simultaneously, we optimize the hand and object poses to achieve accurate object-pose estimations. We then refine physics-based rendering parameters - including roughness, specularity, albedo, hand visibility, skin color reflections, and environmental illumination - to produce precise albedo, and accurate hand illumination and shadow regions. Our approach surpasses state-of-the-art methods in texture reconstruction and, to the best of our knowledge, is the first to account for hand-object interactions in object texture reconstruction.

Keywords

Cite

@article{arxiv.2501.03525,
  title  = {TexHOI: Reconstructing Textures of 3D Unknown Objects in Monocular Hand-Object Interaction Scenes},
  author = {Alakh Aggarwal and Ningna Wang and Xiaohu Guo},
  journal= {arXiv preprint arXiv:2501.03525},
  year   = {2025}
}

Comments

This paper was accepted at CVM 2025 and will appear in IEEE TVCG. This work has been submitted to the IEEE for possible publication

R2 v1 2026-06-28T20:58:21.546Z