中文
相关论文

相关论文: FIction: 4D Future Interaction Prediction from Vid…

200 篇论文

We propose a method for incorporating object interaction and human body dynamics into the task of 3D ego-pose estimation using a head-mounted camera. We use a kinematics model of the human body to represent the entire range of human motion,…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Zhengyi Luo , Ryo Hachiuma , Ye Yuan , Shun Iwase , Kris M. Kitani

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Evonne Ng , Donglai Xiang , Hanbyul Joo , Kristen Grauman

To build general robotic agents that can operate in many environments, it is often imperative for the robot to collect experience in the real world. However, this is often not feasible due to safety, time, and hardware restrictions. We thus…

机器人学 · 计算机科学 2022-12-09 Kenneth Shaw , Shikhar Bahl , Deepak Pathak

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor modeling, where the…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Wenqi Jia , Bolin Lai , Miao Liu , Danfei Xu , James M. Rehg

In order to interact with objects in our environment, humans rely on an understanding of the actions that can be performed on them, as well as their properties. When considering concrete motor actions, this knowledge has been called the…

计算与语言 · 计算机科学 2021-05-12 Ka Chun Lam , Francisco Pereira , Maryam Vaziri-Pashkam , Kristin Woodard , Emalie McMahon

Human perception involves decomposing complex multi-object scenes into time-static object appearance (i.e., size, shape, color) and time-varying object motion (i.e., position, velocity, acceleration). For machines to achieve human-like…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Yeon-Ji Song , Jaein Kim , Suhyung Choi , Jin-Hwa Kim , Byoung-Tak Zhang

In this paper, we present an end-to-end future-prediction model that focuses on pedestrian safety. Specifically, our model uses previous video frames, recorded from the perspective of the vehicle, to predict if a pedestrian will cross in…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Mohamed Chaabane , Ameni Trabelsi , Nathaniel Blanchard , Ross Beveridge

When we physically interact with our environment using our hands, we touch objects and force them to move: contact and motion are defining properties of manipulation. In this paper, we present an active, bottom-up method for the detection…

计算机视觉与模式识别 · 计算机科学 2019-02-05 Konstantinos Zampogiannis , Kanishka Ganguly , Cornelia Fermuller , Yiannis Aloimonos

Anticipating future actions is a key component of intelligence, specifically when it applies to real-time systems, such as robots or autonomous cars. While recent works have addressed prediction of raw RGB pixel values, we focus on…

计算机视觉与模式识别 · 计算机科学 2018-02-09 Fahimeh Rezazadegan , Sareh Shirazi , Mahsa Baktashmotlagh , Larry S. Davis

Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that this should remain…

计算机视觉与模式识别 · 计算机科学 2021-06-18 Paul Henderson , Christoph H. Lampert , Bernd Bickel

We present an approach for pixel-level future prediction given an input image of a scene. We observe that a scene is comprised of distinct entities that undergo motion and present an approach that operationalizes this insight. We implicitly…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Yufei Ye , Maneesh Singh , Abhinav Gupta , Shubham Tulsiani

Video prediction is a crucial task for intelligent agents such as robots and autonomous vehicles, since it enables them to anticipate and act early on time-critical incidents. State-of-the-art video prediction methods typically model the…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Eliyas Suleyman , Paul Henderson , Nicolas Pugeault

We present HOIMotion - a novel approach for human motion forecasting during human-object interactions that integrates information about past body poses and egocentric 3D object bounding boxes. Human motion forecasting is important in many…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Zhiming Hu , Zheming Yin , Daniel Haeufle , Syn Schmitt , Andreas Bulling

In this work we introduce a fully end-to-end approach for action detection in videos that learns to directly predict the temporal bounds of actions. Our intuition is that the process of detecting actions is naturally one of observation and…

计算机视觉与模式识别 · 计算机科学 2017-03-14 Serena Yeung , Olga Russakovsky , Greg Mori , Li Fei-Fei

This paper presents a novel method to predict future human activities from partially observed RGB-D videos. Human activity prediction is generally difficult due to its non-Markovian property and the rich context between human and…

计算机视觉与模式识别 · 计算机科学 2017-08-04 Siyuan Qi , Siyuan Huang , Ping Wei , Song-Chun Zhu

Close human-robot cooperation is a key enabler for new developments in advanced manufacturing and assistive applications. Close cooperation require robots that can predict human actions and intent, and understand human non-verbal cues.…

人机交互 · 计算机科学 2019-02-19 Paul Schydlo , Mirko Rakovic , Lorenzo Jamone , José Santos-Victor

How do humans recognize the action "opening a book" ? We argue that there are two important cues: modeling temporal shape dynamics and modeling functional relationships between humans and objects. In this paper, we propose to represent…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Xiaolong Wang , Abhinav Gupta

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

In this report, we present a cross-view multi-modal object segmentation approach for the object correspondence task in the Ego-Exo4D Correspondence Challenges 2025. Given object queries from one perspective (e.g., ego view), the goal is to…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yuqian Fu , Runze Wang , Yanwei Fu , Danda Pani Paudel , Luc Van Gool

We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models. Given an input video with annotated action periods, the LTA task aims to predict possible future actions. We…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Daoji Huang , Otmar Hilliges , Luc Van Gool , Xi Wang