English
Related papers

Related papers: Building Spatio-temporal Transformers for Egocentr…

200 papers

Transformer-based methods have recently achieved significant success in 3D human pose estimation, owing to their strong ability to model long-range dependencies. However, relying solely on the global attention mechanism is insufficient for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Guangsheng Xu , Guoyi Zhang , Lejia Ye , Shuwei Gan , Xiaohu Zhang , Xia Yang

Augmented reality (AR) displays become more and more popular recently, because of its high intuitiveness for humans and high-quality head-mounted display have rapidly developed. To achieve such displays with augmented information, highly…

Computer Vision and Pattern Recognition · Computer Science 2015-06-22 Kuan-Wen Chen , Chun-Hsin Wang , Xiao Wei , Qiao Liang , Ming-Hsuan Yang , Chu-Song Chen , Yi-Ping Hung

Crucial to the success of training a depth-based 3D hand pose estimator (HPE) is the availability of comprehensive datasets covering diverse camera perspectives, shapes, and pose variations. However, collecting such annotated datasets is…

Computer Vision and Pattern Recognition · Computer Science 2018-05-14 Seungryul Baek , Kwang In Kim , Tae-Kyun Kim

We present EgoRenderer, a system for rendering full-body neural avatars of a person captured by a wearable, egocentric fisheye camera that is mounted on a cap or a VR headset. Our system renders photorealistic novel views of the actor and…

Computer Vision and Pattern Recognition · Computer Science 2021-11-25 Tao Hu , Kripasindhu Sarkar , Lingjie Liu , Matthias Zwicker , Christian Theobalt

Training accurate 3D human pose estimators requires large amount of 3D ground-truth data which is costly to collect. Various weakly or self supervised pose estimation methods have been proposed due to lack of 3D data. Nevertheless, these…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Muhammed Kocabas , Salih Karagoz , Emre Akbas

As a fundamental technique that concerns several vision tasks such as image parsing, action recognition and clothing retrieval, human pose estimation (HPE) has been extensively investigated in recent years. To achieve accurate and reliable…

Computer Vision and Pattern Recognition · Computer Science 2014-11-18 Weipeng Zhang , Jie Shen , Guangcan Liu , Yong Yu

In this paper, we address the problem of estimating the hand pose from the egocentric view when the hand is interacting with objects. Specifically, we propose a method to label a dataset Ego-Siam which contains the egocentric images…

Computer Vision and Pattern Recognition · Computer Science 2021-10-01 Yao Lu , Walterio W. Mayol-Cuevas

Egocentric pose estimation is a fundamental capability for multi-robot collaborative perception in connected autonomy, such as connected autonomous vehicles. During multi-robot operations, a robot needs to know the relative pose between…

Robotics · Computer Science 2025-07-21 Hong Huang , Dongkuan Xu , Hao Zhang , Peng Gao

Estimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Jogendra Nath Kundu , Siddharth Seth , Rahul M , Mugalodi Rakesh , R. Venkatesh Babu , Anirban Chakraborty

Estimating 3D human poses from a monocular video is still a challenging task. Many existing methods' performance drops when the target person is occluded by other objects, or the motion is too fast/slow relative to the scale and speed of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Cheng Yu , Bo Wang , Bo Yang , Robby T. Tan

Occlusions remain one of the key challenges in 3D body pose estimation from single-camera video sequences. Temporal consistency has been extensively used to mitigate their impact but the existing algorithms in the literature do not…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Soumava Kumar Roy , Ilia Badanin , Sina Honari , Pascal Fua

Recovering temporally consistent 3D human body pose, shape and motion from a monocular video is a challenging task due to (self-)occlusions, poor lighting conditions, complex articulated body poses, depth ambiguity, and limited availability…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Sushovan Chanda , Amogh Tiwari , Lokender Tiwari , Brojeshwar Bhowmick , Avinash Sharma , Hrishav Barua

In the field of 3D Human Pose Estimation from monocular videos, the presence of diverse occlusion types presents a formidable challenge. Prior research has made progress by harnessing spatial and temporal cues to infer 3D poses from 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Mehwish Ghafoor , Arif Mahmood , Muhammad Bilal

Although many approaches for multi-human pose estimation in videos have shown profound results, they require densely annotated data which entails excessive man labor. Furthermore, there exists occlusion and motion blur that inevitably lead…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Kyung-Min Jin , Gun-Hee Lee , Seong-Whan Lee

In this report, we present the transferring pretrained video mask autoencoders(VideoMAE) to egocentric tasks for Ego4d Looking at me Challenge. VideoMAE is the data-efficient pretraining model for self-supervised video pre-training and can…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Yinan He , Guo Chen

Human pose estimation focuses on predicting body keypoints to analyze human motion. Currently, most pose estimation tasks rely on conventional RGB cameras. In contrast, event cameras provide high temporal resolution and low latency,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Haoxian Zhou , Chuanzhi Xu , Langyi Chen , Pengfei Ye , Haodong Chen , Yuk Ying Chung , Qiang Qu

In this paper we present our winning entry at the 2018 ECCV PoseTrack Challenge on 3D human pose estimation. Using a fully-convolutional backbone architecture, we obtain volumetric heatmaps per body joint, which we convert to coordinates…

Computer Vision and Pattern Recognition · Computer Science 2019-01-14 István Sárándi , Timm Linder , Kai O. Arras , Bastian Leibe

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

Despite the impressive performance of vision-based pose estimators, they generally fail to perform well under adverse vision conditions and often don't satisfy the privacy demands of customers. As a result, researchers have begun to study…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Vandad Davoodnia , Ali Etemad

Human pose estimation in low-resolution videos presents a fundamental challenge in computer vision. Conventional methods either assume high-quality inputs or employ computationally expensive cascaded processing, which limits their…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Yucheng Jin , Jinyan Chen , Ziyue He , Baojun Han , Furan An