English

GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction

Computer Vision and Pattern Recognition 2026-05-19 v1

Abstract

Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can explicitly provide spatial cues, raw sensor-captured depth maps are extensively noisy and incomplete, limiting their usefulness for fine-grained hand reconstruction. To bridge this gap, we propose GeoHand, a novel framework that unlocks high-quality geometric priors from a frozen foundational monocular geometry estimator (MoGe2). Recognizing that these priors are oriented toward general scenes, we introduce a map-level GeoAdapter to recalibrate the spatial features, specifically adapting them for detailed hand reconstruction. Furthermore, to systematically integrate these adapted priors without overwhelming intrinsic RGB appearance cues, we employ a gated cross-modal token fusion strategy. Finally, to secure precise local articulation, we design a Keypoint-Queried Iterative Refiner (KQIR) that uses projected joint locations to query geometry-aware image features for spatial correction. By combining global geometric disambiguation with local refinement in a unified pipeline, GeoHand achieves state-of-the-art performance on FreiHAND, DexYCB, and HO3Dv3, especially under severe occlusions and hand-object interactions.

Keywords

Cite

@article{arxiv.2605.17354,
  title  = {GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction},
  author = {Weiquan Lin and Yaoqing Hu and Liangchen Dai and Xu Tang and Xingyu Chen},
  journal= {arXiv preprint arXiv:2605.17354},
  year   = {2026}
}
R2 v1 2026-07-22T07:17:15.296Z