English

UmeTrack: Unified multi-view end-to-end hand tracking for VR

Computer Vision and Pattern Recognition 2022-11-02 v1

Abstract

Real-time tracking of 3D hand pose in world space is a challenging problem and plays an important role in VR interaction. Existing work in this space are limited to either producing root-relative (versus world space) 3D pose or rely on multiple stages such as generating heatmaps and kinematic optimization to obtain 3D pose. Moreover, the typical VR scenario, which involves multi-view tracking from wide \ac{fov} cameras is seldom addressed by these methods. In this paper, we present a unified end-to-end differentiable framework for multi-view, multi-frame hand tracking that directly predicts 3D hand pose in world space. We demonstrate the benefits of end-to-end differentiabilty by extending our framework with downstream tasks such as jitter reduction and pinch prediction. To demonstrate the efficacy of our model, we further present a new large-scale egocentric hand pose dataset that consists of both real and synthetic data. Experiments show that our system trained on this dataset handles various challenging interactive motions, and has been successfully applied to real-time VR applications.

Keywords

Cite

@article{arxiv.2211.00099,
  title  = {UmeTrack: Unified multi-view end-to-end hand tracking for VR},
  author = {Shangchen Han and Po-chen Wu and Yubo Zhang and Beibei Liu and Linguang Zhang and Zheng Wang and Weiguang Si and Peizhao Zhang and Yujun Cai and Tomas Hodan and Randi Cabezas and Luan Tran and Muzaffer Akbay and Tsz-Ho Yu and Cem Keskin and Robert Wang},
  journal= {arXiv preprint arXiv:2211.00099},
  year   = {2022}
}

Comments

SIGGRAPH Asia 2022 Conference Papers, 8 pages

R2 v1 2026-06-28T04:53:14.419Z