English
Related papers

Related papers: Guided Attention for Next Active Object @ EGO4D ST…

200 papers

In this paper we introduce a Transformer-based approach to video object segmentation (VOS). To address compounding error and scalability issues of prior work, we propose a scalable, end-to-end method for VOS called Sparse Spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Brendan Duke , Abdalla Ahmed , Christian Wolf , Parham Aarabi , Graham W. Taylor

Egocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to anticipate actions…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Antonino Furnari , Giovanni Maria Farinella

The goal of the Step Grounding task is to locate temporal boundaries of activities based on natural language descriptions. This technical report introduces a Bayesian-VSLNet to address the challenge of identifying such temporal segments in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Carlos Plou , Lorenzo Mur-Labadia , Ruben Martinez-Cantin , Ana C. Murillo

Pure vision transformer architectures are highly effective for short video classification and action recognition tasks. However, due to the quadratic complexity of self attention and lack of inductive bias, transformers are resource…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Edward Fish , Jon Weinbren , Andrew Gilbert

Learning robotic manipulation policies through supervised learning from demonstrations remains challenging when policies encounter execution variations not explicitly covered during training. While incorporating historical context through…

Robotics · Computer Science 2026-03-10 Giovanni Minelli , Giulio Turrisi , Victor Barasuol , Claudio Semini

Previous works for LiDAR-based 3D object detection mainly focus on the single-frame paradigm. In this paper, we propose to detect 3D objects by exploiting temporal information in multiple frames, i.e., the point cloud videos. We empirically…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Junbo Yin , Jianbing Shen , Xin Gao , David Crandall , Ruigang Yang

Utilizing well-trained representations in transfer learning often results in superior performance and faster convergence compared to training from scratch. However, even if such good representations are transferred, a model can easily…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 SeokHyun Seo , Jinwoo Hong , JungWoo Chae , Kyungyul Kim , Sangheum Hwang

Transparent object perception remains a major challenge in computer vision research, as transparency confounds both depth estimation and semantic segmentation. Recent work has explored multi-task learning frameworks to improve robustness,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Gbenga Omotara , Ramy Farag , Seyed Mohamad Ali Tousi , G. N. DeSouza

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Sungjune Park , Hongda Mao , Qingshuang Chen , Yong Man Ro , Yelin Kim

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Spatio-temporal information is key to resolve occlusion and depth ambiguity in 3D pose estimation. Previous methods have focused on either temporal contexts or local-to-global architectures that embed fixed-length spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Junfa Liu , Juan Rojas , Zhijun Liang , Yihui Li , Yisheng Guan

In recent years, we have witnessed an explosive growth of data. Much of this data is video data generated by security cameras, smartphones, and dash cams. The timely analysis of such data is of great practical importance for many emerging…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-06-30 Jayden King , Young Choon Lee

Spatio-temporal action detection (STAD) aims to classify the actions present in a video and localize them in space and time. It has become a particularly active area of research in computer vision because of its explosively emerging…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Peng Wang , Fanwei Zeng , Yuntao Qian

The attention mechanism provides a sequential prediction framework for learning spatial models with enhanced implicit temporal consistency. In this work, we show a systematic design (from 2D to 3D) for how conventional networks and other…

Computer Vision and Pattern Recognition · Computer Science 2021-03-05 Ruixu Liu , Ju Shen , He Wang , Chen Chen , Sen-ching Cheung , Vijayan K. Asari

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Vishakha Lall , Yisi Liu

Triggered by the success of transformers in various visual tasks, the spatial self-attention mechanism has recently attracted more and more attention in the computer vision community. However, we empirically found that a typical vision…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Jiayin Sun , Hong Wang , Qiulei Dong

Long-term action anticipation has become an important task for many applications such as autonomous driving and human-robot interaction. Unlike short-term anticipation, predicting more actions into the future imposes a real challenge with…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Olga Zatsarynna , Emad Bahrami , Yazan Abu Farha , Gianpiero Francesca , Juergen Gall

Attention is an important cognition process of humans, which helps humans concentrate on critical information during their perception and learning. However, although many machine learning models can remember information of data, they have…

Machine Learning · Computer Science 2019-09-06 Guoqiang Zhong , Xin Lin , Kang Chen , Qingyang Li , Kaizhu Huang

We present a visual localization framework based on novel deep attention aware features for autonomous driving that achieves centimeter level localization accuracy. Conventional approaches to the visual localization problem rely on…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Yao Zhou , Guowei Wan , Shenhua Hou , Li Yu , Gang Wang , Xiaofei Rui , Shiyu Song

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta