中文
相关论文

相关论文: EgoPAT3Dv2: Predicting 3D Action Target from 2D Eg…

200 篇论文

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grained dexterous…

We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal…

人工智能 · 计算机科学 2025-07-10 Shaofei Cai , Zhancun Mu , Anji Liu , Yitao Liang

We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a rigid object's perspective. In a HIT, an object is first static relative to the scene, then…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Zhifan Zhu , Siddhant Bansal , Shashank Tripathi , Dima Damen

This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interaction moves toward more intuitive interfaces, the ability to identify targets of non-verbal…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Lukáš Hajdúch , Viktor Kocur

Color-based two-hand 3D pose estimation in the global coordinate system is essential in many applications. However, there are very few datasets dedicated to this task and no existing dataset supports estimation in a non-laboratory…

计算机视觉与模式识别 · 计算机科学 2022-06-13 Fanqing Lin , Tony Martinez

We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various applications, however, it remains largely under-explored. It has been…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Taeho Kang , Kyungjin Lee , Jinrui Zhang , Youngki Lee

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoCast, a bimodal method for 3D human pose forecasting using…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Maria Escobar , Juanita Puentes , Cristhian Forigua , Jordi Pont-Tuset , Kevis-Kokitsi Maninis , Pablo Arbelaez

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Tushar Nagarajan , Kristen Grauman

Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key challenge. Existing…

机器人学 · 计算机科学 2025-06-05 Zhao-Heng Yin , Sherry Yang , Pieter Abbeel

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

We present V-HPOT, a novel approach for improving the cross-domain performance of 3D hand pose estimation from egocentric images across diverse, unseen domains. State-of-the-art methods demonstrate strong performance when trained and tested…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Wiktor Mucha , Michael Wray , Martin Kampel

Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware…

人机交互 · 计算机科学 2024-03-11 Caiyue Xu , Changming Zhang , Yanmin Zhou , Zhipeng Wang , Ping Lu , Bin He

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Christen Millerdurai , Shaoxiang Wang , Yaxu Xie , Vladislav Golyanik , Didier Stricker , Alain Pagani

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data…

机器人学 · 计算机科学 2026-03-24 Xinhai Sun , Xiang Shi , Menglin Zou , Wenlong Huang

The study of human-robot interaction is fundamental to the design and use of robotics in real-world applications. Robots will need to predict and adapt to the actions of human collaborators in order to achieve good performance and improve…

Intention prediction has become a relevant field of research in Human-Machine and Human-Robot Interaction. Indeed, any artificial system (co)-operating with and along humans, designed to assist and coordinate its actions with a human…

机器人学 · 计算机科学 2025-03-20 Anna Belardinelli

In this paper, we propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The key idea is to leverage high-level features linking first- and third-views in a…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Ameya Dhamanaskar , Mariella Dimiccoli , Enric Corona , Albert Pumarola , Francesc Moreno-Noguer

Robotic mapping systems typically approach building metric-semantic scene representations from the robot's own sensors and cameras. However, these "first person" maps inherit the robot's own limitations due to its embodiment or skillset,…

机器人学 · 计算机科学 2026-03-31 Alan Yu , Yun Chang , Christopher Xie , Luca Carlone

Safe and trustworthy Human Robot Interaction (HRI) requires robots not only to complete tasks but also to regulate impedance and speed according to scene context and human proximity. We present SafeHumanoid, an egocentric vision pipeline…