English

Real-Time ESFP: Estimating, Smoothing, Filtering, and Pose-Mapping

Computer Vision and Pattern Recognition 2025-06-27 v1 Robotics

Abstract

This paper presents ESFP, an end-to-end pipeline that converts monocular RGB video into executable joint trajectories for a low-cost 4-DoF desktop arm. ESFP comprises four sequential modules. (1) Estimating: ROMP lifts each frame to a 24-joint 3-D skeleton. (2) Smoothing: the proposed HPSTM-a sequence-to-sequence Transformer with self-attention-combines long-range temporal context with a differentiable forward-kinematics decoder, enforcing constant bone lengths and anatomical plausibility while jointly predicting joint means and full covariances. (3) Filtering: root-normalized trajectories are variance-weighted according to HPSTM's uncertainty estimates, suppressing residual noise. (4) Pose-Mapping: a geometric retargeting layer transforms shoulder-elbow-wrist triples into the uArm's polar workspace, preserving wrist orientation.

Cite

@article{arxiv.2506.21234,
  title  = {Real-Time ESFP: Estimating, Smoothing, Filtering, and Pose-Mapping},
  author = {Qifei Cui and Yuang Zhou and Ruichen Deng},
  journal= {arXiv preprint arXiv:2506.21234},
  year   = {2025}
}
R2 v1 2026-07-01T03:34:28.082Z