面向拥挤场景视频中准确人体姿态估计的研究
计算机视觉与模式识别
2020-10-22 v2
摘要
拥挤场景下的视频人体姿态估计因遮挡、运动模糊、尺度变化和视角改变等成为难题。已有方法总是未能妥善处理该问题,原因在于(1)缺乏时序信息利用;(2)缺乏拥挤场景训练数据。本文从挖掘时序上下文和收集新数据两方面着手,改进拥挤场景视频中的人体姿态估计。具体地,我们首先遵循自顶向下策略检测行人并对每帧进行单人姿态估计。随后,利用源自光流的时序上下文精炼基于帧的姿态估计。具体而言,对某一帧,我们将之前各帧的历史姿态前向传播、将后续各帧的未来姿态后向传播至当前帧,从而实现视频中稳定且准确的人体姿态估计。此外,我们从互联网挖掘与 HIE 数据集相似场景的新数据,以提升训练集多样性。据此,我们的模型在 HIE 挑战赛测试集的 13 个视频中 7 个上取得最佳性能,平均 w_AP 达 56.33。
引用
@article{arxiv.2010.10008,
title = {Towards Accurate Human Pose Estimation in Videos of Crowded Scenes},
author = {Li Yuan and Shuning Chang and Xuecheng Nie and Ziyuan Huang and Yichen Zhou and Yunpeng Chen and Jiashi Feng and Shuicheng Yan},
journal= {arXiv preprint arXiv:2010.10008},
year = {2020}
}
备注
2nd Place in ACM Multimedia Grand Challenge: Human in Events, Track2: Crowd Pose Estimation in Complex Events. ACM Multimedia 2020. arXiv admin note: substantial text overlap with arXiv:2010.08365, arXiv:2010.10007