English
Related papers

Related papers: Hourglass Tokenizer for Efficient Transformer-Base…

200 papers

Training state-of-the-art models for human pose estimation in videos requires datasets with annotations that are really hard and expensive to obtain. Although transformers have been recently utilized for body pose sequence modeling, related…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Fabien Baradel , Romain Brégier , Thibault Groueix , Philippe Weinzaepfel , Yannis Kalantidis , Grégory Rogez

Human Pose Estimation (HPE) based on RGB images has experienced a rapid development benefiting from deep learning. However, event-based HPE has not been fully studied, which remains great potential for applications in extreme scenes and…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Jiaan Chen , Hao Shi , Yaozu Ye , Kailun Yang , Lei Sun , Kaiwei Wang

Over the past few years, the vision transformer and its various forms have gained significance in human pose estimation. By treating image patches as tokens, transformers can capture global relationships wisely, estimate the keypoint tokens…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Anning Li

This paper proposes a unified framework dubbed Multi-view and Temporal Fusing Transformer (MTF-Transformer) to adaptively handle varying view numbers and video length without camera calibration in 3D Human Pose Estimation (HPE). It consists…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Hui Shuai , Lele Wu , Qingshan Liu

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

Although many approaches for multi-human pose estimation in videos have shown profound results, they require densely annotated data which entails excessive man labor. Furthermore, there exists occlusion and motion blur that inevitably lead…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Kyung-Min Jin , Gun-Hee Lee , Seong-Whan Lee

The widespread application of 3D human pose estimation (HPE) is limited by resource-constrained edge devices, requiring more efficient models. A key approach to enhancing efficiency involves designing networks based on the structural…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jialun Cai , Mengyuan Liu , Hong Liu , Shuheng Zhou , Wenhao Li

Transformer architectures have achieved SOTA performance on the human mesh recovery (HMR) from monocular images. However, the performance gain has come at the cost of substantial memory and computational overhead. A lightweight and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ce Zheng , Xianpeng Liu , Guo-Jun Qi , Chen Chen

Previous video-based human pose estimation methods have shown promising results by leveraging aggregated features of consecutive frames. However, most approaches compromise accuracy to mitigate jitter or do not sufficiently comprehend the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Kyung-Min Jin , Byoung-Sung Lim , Gun-Hee Lee , Tae-Kyung Kang , Seong-Whan Lee

Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Wenhao Li , Hong Liu , Runwei Ding , Mengyuan Liu , Pichao Wang , Wenming Yang

There has been a recent surge of interest in introducing transformers to 3D human pose estimation (HPE) due to their powerful capabilities in modeling long-term dependencies. However, existing transformer-based methods treat body joints as…

Computer Vision and Pattern Recognition · Computer Science 2023-02-16 Han Li , Bowen Shi , Wenrui Dai , Hongwei Zheng , Botao Wang , Yu Sun , Min Guo , Chenlin Li , Junni Zou , Hongkai Xiong

Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single person pose estimation. This design relies on heuristic operations such…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yonghui Yu , Jiahang Cai , Xun Wang , Wenwu Yang

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian

Vision Transformers (ViTs) deliver state-of-the-art accuracy but their quadratic attention cost and redundant computations severely hinder deployment on latency and resource-constrained platforms. Existing pruning approaches treat either…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Mohammad Helal Uddin , Liam Seymour , Sabur Baidya

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and a HoloLens headset for data collection, avoiding the use of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jikai Wang , Qifan Zhang , Yu-Wei Chao , Bowen Wen , Xiaohu Guo , Yu Xiang

Video-based 3D human pose and shape estimations are evaluated by intra-frame accuracy and inter-frame smoothness. Although these two metrics are responsible for different ranges of temporal consistency, existing state-of-the-art methods…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Xiaolong Shen , Zongxin Yang , Xiaohan Wang , Jianxin Ma , Chang Zhou , Yi Yang

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Qitao Zhao , Ce Zheng , Mengyuan Liu , Pichao Wang , Chen Chen

Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Yingying Jiao , Zhigang Wang , Zhenguang Liu , Shaojing Fan , Sifan Wu , Zheqi Wu , Zhuoyue Xu

The current methods of video-based 3D human pose estimation have achieved significant progress.However, they still face pressing challenges, such as the underutilization of spatiotemporal bodystructure features in transformers and the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Yang Liu , Zhiyong Zhang

While recent two-stage many-to-one deep learning models have demonstrated great success in 3D human pose estimation, such models are inefficient ways to detect 3D key points in a sequential video relative to one-shot and many-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 David C. Jeong , Hongji Liu , Saunder Salazar , Jessie Jiang , Christopher A. Kitts