中文
相关论文

相关论文: Predicting 4D Hand Trajectory from Monocular Video…

200 篇论文

Haptic exploration is a key skill for both robots and humans to discriminate and handle unknown objects or to recognize familiar objects. Its active nature is evident in humans who from early on reliably acquire sophisticated sensory-motor…

机器人学 · 计算机科学 2020-01-28 Sascha Fleer , Alexandra Moringen , Roberta L. Klatzky , Helge Ritter

We consider the problem of estimating human pose and trajectory by an aerial robot with a monocular camera in near real time. We present a preliminary solution whose distinguishing feature is a dynamic classifier selection architecture. In…

计算机视觉与模式识别 · 计算机科学 2018-12-18 Asanka G Perera , Yee Wei Law , Javaan Chahl

We present a unified framework for understanding 3D hand and object interactions in raw image sequences from egocentric RGB cameras. Given a single RGB image, our model jointly estimates the 3D hand and object poses, models their…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Bugra Tekin , Federica Bogo , Marc Pollefeys

Spatio-temporal information is key to resolve occlusion and depth ambiguity in 3D pose estimation. Previous methods have focused on either temporal contexts or local-to-global architectures that embed fixed-length spatio-temporal…

计算机视觉与模式识别 · 计算机科学 2020-10-21 Junfa Liu , Juan Rojas , Zhijun Liang , Yihui Li , Yisheng Guan

Our work addresses the problem of egocentric human pose estimation from downwards-facing cameras on head-mounted devices (HMD). This presents a challenging scenario, as parts of the body often fall outside of the image or are occluded.…

计算机视觉与模式识别 · 计算机科学 2024-01-29 Hanz Cuevas-Velasquez , Charlie Hewitt , Sadegh Aliakbarian , Tadas Baltrušaitis

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Binjie Zhang , Mike Zheng Shou

Most of the existing deep learning-based methods for 3D hand and human pose estimation from a single depth map are based on a common framework that takes a 2D depth map and directly regresses the 3D coordinates of keypoints, such as hand or…

计算机视觉与模式识别 · 计算机科学 2018-08-17 Gyeongsik Moon , Ju Yong Chang , Kyoung Mu Lee

Multi-view egocentric hand tracking is a challenging task and plays a critical role in VR interaction. In this report, we present a method that uses multi-view input images and camera extrinsic parameters to estimate both hand shape and…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Minqiang Zou , Zhi Lv , Riqiang Jin , Tian Zhan , Mochen Yu , Yao Tang , Jiajun Liang

We present an online approach to efficiently and simultaneously detect and track the 2D pose of multiple people in a video sequence. We build upon Part Affinity Field (PAF) representation designed for static images, and propose an…

计算机视觉与模式识别 · 计算机科学 2019-06-14 Yaadhav Raaj , Haroon Idrees , Gines Hidalgo , Yaser Sheikh

Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Zhongwei Qiu , Qiansheng Yang , Jian Wang , Dongmei Fu

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Sen Yang , Wen Heng , Gang Liu , Guozhong Luo , Wankou Yang , Gang Yu

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability to predict plausible…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Rustin Soraki , Homanga Bharadhwaj , Ali Farhadi , Roozbeh Mottaghi

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Christen Millerdurai , Shaoxiang Wang , Yaxu Xie , Vladislav Golyanik , Didier Stricker , Alain Pagani

Dynamic view synthesis has seen significant advances, yet reconstructing scenes from uncalibrated, casual video remains challenging due to slow optimization and complex parameter estimation. In this work, we present Instant4D, a monocular…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zhanpeng Luo , Haoxi Ran , Li Lu

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery, HaMeR, follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Georgios Pavlakos , Dandan Shan , Ilija Radosavovic , Angjoo Kanazawa , David Fouhey , Jitendra Malik

Understanding human intentions and actions through egocentric videos is important on the path to embodied artificial intelligence. As a branch of egocentric vision techniques, hand trajectory prediction plays a vital role in comprehending…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Junyi Ma , Xieyuanli Chen , Wentao Bao , Jingyi Xu , Hesheng Wang

We present a new computational model for gaze prediction in egocentric videos by exploring patterns in temporal shift of gaze fixations (attention transition) that are dependent on egocentric manipulation tasks. Our assumption is that the…

计算机视觉与模式识别 · 计算机科学 2018-12-05 Yifei Huang , Minjie Cai , Zhenqiang Li , Yoichi Sato

Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable high-order temporal…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Dingkun Wei , Zehong Shen , Yan Xia , Georgios Pavlakos , Yujun Shen , Xiaowei Zhou

Humans naturally integrate vision and haptics for robust object perception during manipulation. The loss of either modality significantly degrades performance. Inspired by this multisensory integration, prior object pose estimation research…

机器人学 · 计算机科学 2025-09-12 Hongyu Li , Mingxi Jia , Tuluhan Akbulut , Yu Xiang , George Konidaris , Srinath Sridhar

We present a novel appearance-based approach for pose estimation of a human hand using the point clouds provided by the low-cost Microsoft Kinect sensor. Both the free-hand case, in which the hand is isolated from the surrounding…

计算机视觉与模式识别 · 计算机科学 2016-04-08 Pasquale Coscia , Francesco A. N. Palmieri , Francesco Castaldo , Alberto Cavallo