中文
相关论文

相关论文: Building Spatio-temporal Transformers for Egocentr…

200 篇论文

A self-driving perception model aims to extract 3D semantic representations from multiple cameras collectively into the bird's-eye-view (BEV) coordinate frame of the ego car in order to ground downstream planner. Existing perception methods…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Jiachen Lu , Zheyuan Zhou , Xiatian Zhu , Hang Xu , Li Zhang

Reconstructing 3D hand mesh is challenging but an important task for human-computer interaction and AR/VR applications. In particular, RGB and/or depth cameras have been widely used in this task. However, methods using these conventional…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ryosei Hara , Wataru Ikeda , Masashi Hatano , Mariko Isogawa

Current human pose estimation systems focus on retrieving an accurate 3D global estimate of a single person. Therefore, this paper presents one of the first 3D multi-person human pose estimation systems that is able to work in real-time and…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Pawel Knap , Peter Hardy , Alberto Tamajo , Hwasup Lim , Hansung Kim

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jaeyoung Choi , Hyeondong Kim , Yujin Kim , Daehee Park

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Yilin Wen , Hao Pan , Lei Yang , Jia Pan , Taku Komura , Wenping Wang

Using an ego-centric camera to do localization and tracking is highly needed for urban navigation and indoor assistive system when GPS is not available or not accurate enough. The traditional hand-designed feature tracking and estimation…

计算机视觉与模式识别 · 计算机科学 2018-12-04 Liang Yang , Hao Jiang , Jizhong Xiao , Zhouyuan Huo

Hand pose represents key information for action recognition in the egocentric perspective, where the user is interacting with objects. We propose to improve egocentric 3D hand pose estimation based on RGB frames only by using pseudo-depth…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Wiktor Mucha , Michael Wray , Martin Kampel

With the recent advancements in single-image-based human mesh recovery, there is a growing interest in enhancing its performance in certain extreme scenarios, such as occlusion, while maintaining overall model accuracy. Although obtaining…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Wendi Yang , Zihang Jiang , Shang Zhao , S. Kevin Zhou

We introduce HYPERPOSE, a novel 3D human pose estimation framework that performs spatio-temporal reasoning entirely within the Lorentz model of hyperbolic space $\mathbb{H}^d$ to natively preserve the hierarchical tree topology of the human…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Vinduja Thekkath , Ashish Musale , Ajay Waghumbare , Upasna Singh

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the…

计算机视觉与模式识别 · 计算机科学 2023-03-17 Zhongwei Qiu , Yang Qiansheng , Jian Wang , Haocheng Feng , Junyu Han , Errui Ding , Chang Xu , Dongmei Fu , Jingdong Wang

The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the wearer's limbs. While…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Kyungwon Cho , Hanbyul Joo

This technical report introduces our solution, MEEV, proposed to the EgoBody Challenge at ECCV 2022. Captured from head-mounted devices, the dataset consists of human body shape and motion of interacting people. The EgoBody dataset has…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Nicolas Monet , Dongyoon Wee

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Shuo Liang , Yiwu Zhong , Zi-Yuan Hu , Yeyao Tao , Liwei Wang

Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Wentao Bao , Lele Chen , Libing Zeng , Zhong Li , Yi Xu , Junsong Yuan , Yu Kong

Camera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving…

计算机视觉与模式识别 · 计算机科学 2020-04-10 Jogendra Nath Kundu , Siddharth Seth , Varun Jampani , Mugalodi Rakesh , R. Venkatesh Babu , Anirban Chakraborty

Vision-based ego-lane inference using High-Definition (HD) maps is essential in autonomous driving and advanced driver assistance systems. The traditional approach necessitates well-calibrated cameras, which confines variation of camera…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Chaehyeon Song , Sungho Yoon , Minhyeok Heo , Ayoung Kim , Sujung Kim

The success or failure of modern computer-assisted surgery procedures hinges on the precise six-degree-of-freedom (6DoF) position and orientation (pose) estimation of tracked instruments and tissue. In this paper, we present HMD-EgoPose, a…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Mitchell Doughty , Nilesh R. Ghugre

The emergence of data-driven approaches for control and planning in robotics have highlighted the need for developing experimental robotic platforms for data collection. However, their implementation is often complex and expensive, in…

机器人学 · 计算机科学 2022-02-02 Quentin Possamaï , Steeven Janny , Guillaume Bono , Madiha Nadri , Laurent Bako , Christian Wolf

Supervised approaches to 3D pose estimation from single images are remarkably effective when labeled data is abundant. However, as the acquisition of ground-truth 3D labels is labor intensive and time consuming, recent attention has shifted…

计算机视觉与模式识别 · 计算机科学 2022-06-30 Soumava Kumar Roy , Leonardo Citraro , Sina Honari , Pascal Fua

Humans have an innate ability to sense their surroundings, as they can extract the spatial representation from the egocentric perception and form an allocentric semantic map via spatial transformation and memory updating. However, endowing…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Chang Chen , Jiaming Zhang , Kailun Yang , Kunyu Peng , Rainer Stiefelhagen