中文
相关论文

相关论文: Intention-Conditioned Long-Term Human Egocentric A…

200 篇论文

Egocentric action anticipation is the task of predicting the future actions a camera wearer will likely perform based on past video observations. While in a real-world system it is fundamental to output such predictions before the action…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Antonino Furnari , Giovanni Maria Farinella

Video anticipation is the task of predicting one/multiple future representation(s) given limited, partial observation. This is a challenging task due to the fact that given limited observation, the future representation can be highly…

计算机视觉与模式识别 · 计算机科学 2020-10-12 Sadegh Aliakbarian

While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple navigational instructions. Transitioning to genuinely…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Huan Zheng , Yucheng Zhou , Tianyi Yan , Jiayi Su , Hongjun Chen , Dubing Chen , Xingtai Gui , Wencheng Han , Runzhou Tao , Zhongying Qiu , Jianfei Yang , Jianbing Shen

Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works have explored leveraging human data to alleviate this…

机器人学 · 计算机科学 2026-05-01 Chengyang Li , Kaiyi Xiong , Yuan Xu , Lei Qian , Yizhou Wang , Wentao Zhu

We introduce a framework that predicts the goals behind observable human action in video. Motivated by evidence in developmental psychology, we leverage video of unintentional action to learn video representations of goals without direct…

计算机视觉与模式识别 · 计算机科学 2020-12-17 Dave Epstein , Carl Vondrick

Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model interactions by…

计算机视觉与模式识别 · 计算机科学 2024-01-12 Debaditya Roy , Ramanathan Rajendiran , Basura Fernando

During collaborative tasks, human behavior is guided by multiple levels of intentions that evolve over time, such as task sequence preferences and interaction strategies. To adapt to these changing preferences and promptly correct any…

机器人学 · 计算机科学 2025-06-18 Zhe Huang , Ye-Ji Mun , Fatemeh Cheraghi Pouria , Katherine Driggs-Campbell

In this paper, we address the problem of short-term action anticipation, i.e., we want to predict an upcoming action one second before it happens. We propose to harness high-level intent information to anticipate actions that will take…

计算机视觉与模式识别 · 计算机科学 2023-06-28 Olga Zatsarynna , Juergen Gall

We present a generative approach to forecast long-term future human behavior in 3D, requiring only weak supervision from readily available 2D human action data. This is a fundamental task enabling many downstream applications. The required…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Christian Diller , Thomas Funkhouser , Angela Dai

Human behavior has the nature of mutual dependencies, which requires human-robot interactive systems to predict surrounding agents trajectories by modeling complex social interactions, avoiding collisions and executing safe path planning.…

机器人学 · 计算机科学 2026-04-21 Yuxiang Zhao , Wei Huang , Haipeng Zeng , Huan Zhao , Yujie Song

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To incorporate the…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Evangelos Kazakos , Jaesung Huh , Arsha Nagrani , Andrew Zisserman , Dima Damen

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interaction with social…

机器人学 · 计算机科学 2025-04-09 Hassan Ali , Philipp Allgeuer , Stefan Wermter

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Sungjune Park , Hongda Mao , Qingshuang Chen , Yong Man Ro , Yelin Kim

In collaborative human-robot manipulation, a robot must predict human intents and adapt its actions accordingly to smoothly execute tasks. However, the human's intent in turn depends on actions the robot takes, creating a chicken-or-egg…

机器人学 · 计算机科学 2024-06-04 Kushal Kedia , Atiksh Bhardwaj , Prithwish Dan , Sanjiban Choudhury

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

机器学习 · 计算机科学 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Yiming Li , Ziang Cao , Andrew Liang , Benjamin Liang , Luoyao Chen , Hang Zhao , Chen Feng

Fluent and safe interactions of humans and robots require both partners to anticipate the others' actions. A common approach to human intention inference is to model specific trajectories towards known goals with supervised classifiers.…

机器人学 · 计算机科学 2017-02-28 Judith Bütepage , Hedvig Kjellström , Danica Kragic

We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models. Given an input video with annotated action periods, the LTA task aims to predict possible future actions. We…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Daoji Huang , Otmar Hilliges , Luc Van Gool , Xi Wang

Predicting future human behavior from an input human video is a useful task for applications such as autonomous driving and robotics. While most previous works predict a single future, multiple futures with different behavior can…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Naoya Fushishita , Antonio Tejero-de-Pablos , Yusuke Mukuta , Tatsuya Harada

In this paper, we tackle the problem of egocentric action anticipation, i.e., predicting what actions the camera wearer will perform in the near future and which objects they will interact with. Specifically, we contribute Rolling-Unrolling…

计算机视觉与模式识别 · 计算机科学 2020-05-11 Antonino Furnari , Giovanni Maria Farinella