中文
相关论文

相关论文: Ego-centric Predictive Model Conditioned on Hand T…

200 篇论文

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

计算机视觉与模式识别 · 计算机科学 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

Dexterous manipulation is essential for real-world robot autonomy, mirroring the central role of human hand coordination in daily activity. Humans rely on rich multimodal perception--vision, sound, and language-guided intent--to perform…

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However,…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Qiaohui Chu , Haoyu Zhang , Meng Liu , Yisen Feng , Haoxiang Shi , Liqiang Nie

Future activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a…

计算机视觉与模式识别 · 计算机科学 2021-11-24 Zhaobo Qi , Shuhui Wang , Chi Su , Li Su , Qingming Huang , Qi Tian

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans'…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Danny Tran , Roberto Martín-Martín , Kristen Grauman

Egocentric action anticipation aims to predict the future actions the camera wearer will perform from the observation of the past. While predictions about the future should be available before the predicted events take place, most…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Antonino Furnari , Giovanni Maria Farinella

Forecasting future events based on evidence of current conditions is an innate skill of human beings, and key for predicting the outcome of any decision making. In artificial vision for example, we would like to predict the next human…

计算机视觉与模式识别 · 计算机科学 2022-06-03 Tsung-Ming Tai , Giuseppe Fiameni , Cheng-Kuang Lee , Simon See , Oswald Lanz

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretraining, leaving physical dynamics to be learned from…

机器人学 · 计算机科学 2026-03-24 Teli Ma , Jia Zheng , Zifan Wang , Chunli Jiang , Andy Cui , Junwei Liang , Shuo Yang

Can we turn a video prediction model into a robot policy? Videos, including those of humans or teleoperated robots, capture rich physical interactions. However, most of them lack labeled actions, which limits their use in robot learning. We…

机器人学 · 计算机科学 2026-03-31 Sandeep Routray , Hengkai Pan , Unnat Jain , Shikhar Bahl , Deepak Pathak

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented egocentric videos…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Ye Yuan , Kris Kitani

In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Wenqi Jia , Miao Liu , Hao Jiang , Ishwarya Ananthabhotla , James M. Rehg , Vamsi Krishna Ithapu , Ruohan Gao

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Zihui Xue , Kristen Grauman

Egocentric action anticipation consists in predicting a future action the camera wearer will perform from egocentric video. While the task has recently attracted the attention of the research community, current approaches assume that the…

计算机视觉与模式识别 · 计算机科学 2022-02-10 Ivan Rodin , Antonino Furnari , Dimitrios Mavroeidis , Giovanni Maria Farinella

Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Eadom Dessalene , Chinmaya Devaraj , Michael Maynord , Cornelia Fermuller , Yiannis Aloimonos

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial…

机器人学 · 计算机科学 2026-05-22 Wenxuan Guo , Ziyuan Li , Meng Zhang , Yichen Liu , Yimeng Dong , Chuxi Xu , Yunfei Wei , Ze Chen , Erjin Zhou , Jianjiang Feng

Robotic manipulation involves kinematic and semantic transitions that are inherently coupled via underlying actions. However, existing approaches plan within either semantic or latent space without explicitly aligning these cross-modal…

机器人学 · 计算机科学 2026-04-01 Andrew Jeong , Jaemin Kim , Sebin Lee , Sung-Eui Yoon

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Zi-Yi Dou , Xitong Yang , Tushar Nagarajan , Huiyu Wang , Jing Huang , Nanyun Peng , Kris Kitani , Fu-Jen Chu