中文

基于单目 RGB 视频 3D landmarks 的物理信息人体姿态估计

计算机视觉与模式识别 2025-12-09 v1

摘要

提供自动化体能训练的应用日益受到欢迎,例如体育康复。这些应用依赖于使用单目视频流进行 accurate and robust pose estimation。像 BlazePose 这样的 state-of-the-art 模型在实时姿态跟踪方面表现出色,但其 lack of解剖约束 indicates 可通过包含物理知识来实现改进。我们提出一种 real-time post-processing algorithm,融合 BlazePose 3D 和 2D 估计的优势,采用加权 optimization,惩罚偏离预期 bone length 和生物力学 model 的 deviation。骨骼长度估计被 refined 到 individual anatomy,使用适应 measurement trust 的 Kalman filter。使用 Physio2.2M 数据集进行评估,结果显示 3D MPJPE 减少了 10.2%,体 segments 之间角度 error 减少了 16.6%。该方法提供了一种基于计算效率高的 video-to-3D pose estimation 的 robust、解剖一致的人体姿态估计,适用于 consumer-level laptop 和 mobile 设备上的自动化体育康复、健康和 sports coaching。该 refinement 在后端运行,仅使用 anonymized data。

关键词

引用

@article{arxiv.2512.06783,
  title  = {Physics Informed Human Posture Estimation Based on 3D Landmarks from Monocular RGB-Videos},
  author = {Tobias Leuthold and Michele Xiloyannis and Yves Zimmermann},
  journal= {arXiv preprint arXiv:2512.06783},
  year   = {2025}
}

备注

16 pages, 5 figures