English

Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation

Computer Vision and Pattern Recognition 2026-04-08 v3

Abstract

Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available and show that gesture semantics can serve as a powerful inductive bias for 3D pose estimation. We present a two-stage framework: gesture-aware pretraining that learns an informative embedding space using coarse and fine gesture labels from InterHand2.6M, followed by a per-joint token Transformer guided by gesture embeddings as intermediate representations for final regression of MANO hand parameters. Training is driven by a layered objective over parameters, joints, and structural constraints. Experiments on InterHand2.6M demonstrate that gesture-aware pretraining consistently improves single-hand accuracy over the state-of-the-art EANet baseline, and that the benefit transfers across architectures without any modification.

Keywords

Cite

@article{arxiv.2603.17396,
  title  = {Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation},
  author = {Rui Hong and Jana Kosecka},
  journal= {arXiv preprint arXiv:2603.17396},
  year   = {2026}
}

Comments

6 pages, 6 figures

R2 v1 2026-07-01T11:25:37.062Z