中文
相关论文

相关论文: Predicting 4D Hand Trajectory from Monocular Video…

200 篇论文

3D hand pose estimation and shape recovery are challenging tasks in computer vision. We introduce a novel framework HandTailor, which combines a learning-based hand module and an optimization-based tailor module to achieve high-precision…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Jun Lv , Wenqiang Xu , Lixin Yang , Sucheng Qian , Chongzhao Mao , Cewu Lu

Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within the environment remains a challenge. Existing methods primarily focus on relative motion…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Hiroyuki Deguchi , Ryosuke Hori , Kotaro Amaya , Tsubasa Maruyama , Mitsunori Tada , Hideo Saito

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yinqiao Wang , Hao Xu , Pheng-Ann Heng , Chi-Wing Fu

The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the wearer's limbs. While…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Kyungwon Cho , Hanbyul Joo

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Evonne Ng , Shiry Ginosar , Trevor Darrell , Hanbyul Joo

The number of static human poses is limited, it is hard to retrieve the exact videos using one single pose as the clue. However, with a pose sequence or a dynamic gesture as the keyword, retrieving specific videos becomes more feasible. We…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Cheng Zhang

We consider predicting the user's head motion in 360-degree videos, with 2 modalities only: the past user's positions and the video content (not knowing other users' traces). We make two main contributions. First, we re-examine existing…

计算机视觉与模式识别 · 计算机科学 2021-04-15 Miguel Fabian Romero Rondon , Lucile Sassatelli , Ramon Aparicio Pardo , Frederic Precioso

Embodied world models aim to predict and interact with the physical world through visual observations and actions. However, existing models struggle to accurately translate low-level actions (e.g., joint positions) into precise robotic…

机器人学 · 计算机科学 2026-04-01 Taiyi Su , Jian Zhu , Yaxuan Li , Chong Ma , Jianjun Zhang , Zitai Huang , Hanli Wang , Yi Xu

Soft robotic hand shows considerable promise for various grasping applications. However, the sensing and reconstruction of the robot pose will cause limitation during the design and fabrication. In this work, we present a novel 3D pose…

机器人学 · 计算机科学 2023-08-08 Haihang Wang , He Xu , Yihan Meng

We propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Haixin Shi , Yinlin Hu , Daniel Koguciuk , Juan-Ting Lin , Mathieu Salzmann , David Ferstl

This paper presents a novel 3D human pose estimation approach using a single stream of asynchronous events as input. Most of the state-of-the-art approaches solve this task with RGB cameras, however struggling when subjects are moving fast.…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Gianluca Scarpellini , Pietro Morerio , Alessio Del Bue

Hand pose estimation from a single image has many applications. However, approaches to full 3D body pose estimation are typically trained on day-to-day activities or actions. As such, detailed hand-to-hand interactions are poorly…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Video-based human pose estimation remains challenged by motion blur, occlusion, and complex spatiotemporal dynamics. Existing methods often rely on heatmaps or implicit spatio-temporal feature aggregation, which limits joint topology…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Quang Dang Huynh , Xuefei Yin , Andrew Busch , Hugo G. Espinosa , Alan Wee-Chung Liew , Matthew T. O. Worsey , Yanming Zhu

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Ye Yuan , Umar Iqbal , Pavlo Molchanov , Kris Kitani , Jan Kautz

Video transformers have recently emerged as an effective alternative to convolutional networks for action classification. However, most prior video transformers adopt either global space-time attention or hand-defined strategies to compare…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Jue Wang , Lorenzo Torresani

We introduce Mono4DGS-HDR, the first system for reconstructing renderable 4D high dynamic range (HDR) scenes from unposed monocular low dynamic range (LDR) videos captured with alternating exposures. To tackle such a challenging problem, we…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Jinfeng Liu , Lingtong Kong , Mi Zhou , Jinwen Chen , Dan Xu

We present an approach for real-time, robust and accurate hand pose estimation from moving egocentric RGB-D cameras in cluttered real environments. Existing methods typically fail for hand-object interactions in cluttered scenes imaged from…

计算机视觉与模式识别 · 计算机科学 2018-11-26 Franziska Mueller , Dushyant Mehta , Oleksandr Sotnychenko , Srinath Sridhar , Dan Casas , Christian Theobalt

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks.…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Qi Xia , Peishan Cong , Ziyi Wang , Yujing Sun , Qin Sun , Xinge Zhu , Mao Ye , Ruigang Yang , Yuexin Ma

Tracking human object interaction from videos is important to understand human behavior from the rapidly growing stream of video data. Previous video-based methods require predefined object templates while single-image-based methods are…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Xianghui Xie , Jan Eric Lenssen , Gerard Pons-Moll