中文
相关论文

相关论文: Leveraging Next-Active Objects for Context-Aware A…

200 篇论文

The introduction of Transformer model has led to tremendous advancements in sequence modeling, especially in text domain. However, the use of attention-based models for video understanding is still relatively unexplored. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2021-03-19 Saurabh Sahu , Palash Goyal

The task of predicting future actions from a video is crucial for a real-world agent interacting with others. When anticipating actions in the distant future, we humans typically consider long-term relations over the whole sequence of…

计算机视觉与模式识别 · 计算机科学 2022-05-30 Dayoung Gong , Joonseok Lee , Manjin Kim , Seong Jong Ha , Minsu Cho

Anticipating how a person will interact with objects in an environment is essential for activity understanding, but existing methods are limited to the 2D space of video frames-capturing physically ungrounded predictions of "what" and…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Kumar Ashutosh , Georgios Pavlakos , Kristen Grauman

In autonomous driving and robotics, there is a growing interest in utilizing short-term historical data to enhance multi-camera 3D object detection, leveraging the continuous and correlated nature of input video streams. Recent work has…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Seokha Moon , Hongbeen Park , Jungphil Kwon , Jaekoo Lee , Jinkyu Kim

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

Predicting human interaction is challenging as the on-going activity has to be inferred based on a partially observed video. Essentially, a good algorithm should effectively model the mutual influence between the two interacting subjects.…

计算机视觉与模式识别 · 计算机科学 2017-05-29 Yichao Yan , Bingbing Ni , Xiaokang Yang

As drone technology advances, using unmanned aerial vehicles for aerial surveys has become the dominant trend in modern low-altitude remote sensing. The surge in aerial video data necessitates accurate prediction for future scenarios and…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Liangyu Xu , Wanxuan Lu , Hongfeng Yu , Yongqiang Mao , Hanbo Bi , Chenglong Liu , Xian Sun , Kun Fu

This paper proposes a method for long-term action anticipation (LTA), the task of predicting action labels and their duration in a video given the observation of an initial untrimmed video interval. We build on an encoder-decoder…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Alberto Maté , Mariella Dimiccoli

Spatio-temporal scene graphs represent interactions in a video by decomposing scenes into individual objects and their pair-wise temporal relationships. Long-term anticipation of the fine-grained pair-wise relationships between objects is a…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Rohith Peddi , Saksham Singh , Saurabh , Parag Singla , Vibhav Gogate

Video-based human-object interaction (HOI) understanding requires both detecting ongoing interactions and anticipating their future evolution. However, existing methods usually treat anticipation as a downstream forecasting task built on…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Yuanhao Luo , Di Wen , Kunyu Peng , Ruiping Liu , Junwei Zheng , Yufan Chen , Jiale Wei , Rainer Stiefelhage

A world model is essential for an agent to predict the future and plan in domains such as autonomous driving and robotics. To achieve this, recent advancements have focused on video generation, which has gained significant attention due to…

人工智能 · 计算机科学 2025-03-13 Youngjoon Jeong , Junha Chun , Soonwoo Cha , Taesup Kim

Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, we find that…

计算机视觉与模式识别 · 计算机科学 2024-02-22 Binglu Wang , Chenxi Guo , Yang Jin , Haisheng Xia , Nian Liu

An effective understanding of the environment and accurate trajectory prediction of surrounding dynamic obstacles are indispensable for intelligent mobile systems (e.g. autonomous vehicles and social robots) to achieve safe and high-quality…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Jiachen Li , Hengbo Ma , Zhihao Zhang , Jinning Li , Masayoshi Tomizuka

Learning an egocentric action recognition model from video data is challenging due to distractors (e.g., irrelevant objects) in the background. Further integrating object information into an action model is hence beneficial. Existing…

计算机视觉与模式识别 · 计算机科学 2022-05-04 Victor Escorcia , Ricardo Guerrero , Xiatian Zhu , Brais Martinez

Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric perspective, which serves as a key bridge between humans and machines in action…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mengmeng Ge , Takashi Isobe , Xu Jia , Yanan Sun , Zetong Yang , Weinong Wang , Dong Zhou , Dong Li , Huchuan Lu , Emad Barsoum

Recognising the characteristics of objects while a robot handles them is crucial for adjusting motions that ensure stable and efficient interactions with containers. Ahead of realising stable and efficient robot motions for…

机器人学 · 计算机科学 2024-03-19 Namiko Saito , Joao Moura , Hiroki Uchida , Sethu Vijayakumar

Visual search is important in our daily life. The efficient allocation of visual attention is critical to effectively complete visual search tasks. Prior research has predominantly modelled the spatial allocation of visual attention in…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yini Fang , Jingling Yu , Haozheng Zhang , Ralf van der Lans , Bertram Shi

In this paper, we tackle the problem of egocentric action anticipation, i.e., predicting what actions the camera wearer will perform in the near future and which objects they will interact with. Specifically, we contribute Rolling-Unrolling…

计算机视觉与模式识别 · 计算机科学 2020-05-11 Antonino Furnari , Giovanni Maria Farinella

We present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-noun structure of action labels in egocentric video datasets.…

计算机视觉与模式识别 · 计算机科学 2021-02-17 Swathikiran Sudhakaran , Sergio Escalera , Oswald Lanz

Proactive robot assistance enables a robot to anticipate and provide for a user's needs without being explicitly asked. We formulate proactive assistance as the problem of the robot anticipating temporal patterns of object movements…

机器人学 · 计算机科学 2023-09-13 Maithili Patel , Sonia Chernova