中文
相关论文

相关论文: Ego-centric Predictive Model Conditioned on Hand T…

200 篇论文

How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we extend the classic hand trajectory prediction task to two tasks…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Chen Bao , Jiarui Xu , Xiaolong Wang , Abhinav Gupta , Homanga Bharadhwaj

Anticipating human actions is an important task that needs to be addressed for the development of reliable intelligent agents, such as self-driving cars or robot assistants. While the ability to make future predictions with high accuracy is…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

Short-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on…

计算机视觉与模式识别 · 计算机科学 2023-06-26 Sanket Thakur , Cigdem Beyan , Pietro Morerio , Vittorio Murino , Alessio Del Bue

Egocentric action anticipation is the task of predicting the future actions a camera wearer will likely perform based on past video observations. While in a real-world system it is fundamental to output such predictions before the action…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Antonino Furnari , Giovanni Maria Farinella

Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination…

机器人学 · 计算机科学 2026-03-31 Qiqi Liu , Huan Xu , Jingyu Li , Bin Sun , Zhihui Hao , Dangen She , Xiatian Zhu , Li Zhang

Learning an agent model that behaves like humans-capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective-is a fundamental challenge in computer vision. Existing methods…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Lu Chen , Yizhou Wang , Shixiang Tang , Qianhong Ma , Tong He , Wanli Ouyang , Xiaowei Zhou , Hujun Bao , Sida Peng

We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models. Given an input video with annotated action periods, the LTA task aims to predict possible future actions. We…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Daoji Huang , Otmar Hilliges , Luc Van Gool , Xi Wang

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Si-En Hong , James Tribble , Alexander Lake , Hao Wang , Chaoyi Zhou , Ashish Bastola , Siyu Huang , Eisa Chaudhary , Brian Canada , Ismahan Arslan-Ari , Abolfazl Razi

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action,…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Haoyu Zhen , Xiaowen Qiu , Peihao Chen , Jincheng Yang , Xin Yan , Yilun Du , Yining Hong , Chuang Gan

Most existing vision-language-action (VLA) models for robotic manipulation lack progress awareness, typically relying on hand-crafted heuristics for task termination. This limitation is particularly severe in long-horizon tasks involving…

机器人学 · 计算机科学 2026-03-31 Hongyu Yan , Qiwei Li , Jiaolong Yang , Yadong Mu

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Chaitanya Patel , Hiroki Nakamura , Yuta Kyuragi , Kazuki Kozuka , Juan Carlos Niebles , Ehsan Adeli

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action…

计算机视觉与模式识别 · 计算机科学 2021-01-20 Yu Wu , Linchao Zhu , Xiaohan Wang , Yi Yang , Fei Wu

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive…

机器人学 · 计算机科学 2025-12-23 Max Argus , Jelena Bratulic , Houman Masnavi , Maxim Velikanov , Nick Heppert , Abhinav Valada , Thomas Brox

Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation often fails to…

机器人学 · 计算机科学 2026-02-27 Daesol Cho , Youngseok Jang , Danfei Xu , Sehoon Ha

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate modality. In…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Miao Liu , Siyu Tang , Yin Li , James Rehg

Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs) are essential to…