English
Related papers

Related papers: Learning to Recognize Actions on Objects in Egocen…

200 papers

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

Language-driven action localization in videos is a challenging task that involves not only visual-linguistic matching but also action boundary prediction. Recent progress has been achieved through aligning language query to video segments,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Shuo Yang , Xinxiao Wu

Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confined in limited types…

Robotics · Computer Science 2026-02-05 Jialiang Li , Yi Qiao , Yunhan Guo , Changwen Chen , Wenzhao Lian

Object-centric representations form the basis of human perception, and enable us to reason about the world and to systematically generalize to new settings. Currently, most works on unsupervised object discovery focus on slot-based…

Machine Learning · Computer Science 2022-11-21 Sindy Löwe , Phillip Lippe , Maja Rudolph , Max Welling

Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentation and tracking in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Yash Bhalgat , Vadim Tschernezki , Iro Laina , João F. Henriques , Andrea Vedaldi , Andrew Zisserman

Facial Action Coding System is an approach for modeling the complexity of human emotional expression. Automatic action unit (AU) detection is a crucial research area in human-computer interaction. This paper describes our submission to the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Duy Le Hoai , Eunchae Lim , Eunbin Choi , Sieun Kim , Sudarshan Pant , Guee-Sang Lee , Soo-Huyng Kim , Hyung-Jeong Yang

Action recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, existing methods usually…

Computer Vision and Pattern Recognition · Computer Science 2022-09-05 Xiaodong Chen , Xinchen Liu , Wu Liu , Kun Liu , Dong Wu , Yongdong Zhang , Tao Mei

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Tushar Nagarajan , Yanghao Li , Christoph Feichtenhofer , Kristen Grauman

We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models. Given an input video with annotated action periods, the LTA task aims to predict possible future actions. We…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Daoji Huang , Otmar Hilliges , Luc Van Gool , Xi Wang

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

Multimedia · Computer Science 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

Understanding the world from multiple perspectives is essential for intelligent systems operating together, where segmenting common objects across different views remains an open problem. We introduce a new approach that re-defines…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Lorenzo Mur-Labadia , Maria Santos-Villafranca , Jesus Bermudez-Cameo , Alejandro Perez-Yus , Ruben Martinez-Cantin , Jose J. Guerrero

We propose a soft attention based model for the task of action recognition in videos. We use multi-layered Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units which are deep both spatially and temporally. Our model…

Machine Learning · Computer Science 2016-02-16 Shikhar Sharma , Ryan Kiros , Ruslan Salakhutdinov

Activity recognition from long unstructured egocentric photo-streams has several applications in assistive technology such as health monitoring and frailty detection, just to name a few. However, one of its main technical challenges is to…

Computer Vision and Pattern Recognition · Computer Science 2017-08-29 Alejandro Cartas , Mariella Dimiccoli , Petia Radeva

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

Computer Vision and Pattern Recognition · Computer Science 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

Contextual information plays an important role in action recognition. Local operations have difficulty to model the relation between two elements with a long-distance interval. However, directly modeling the contextual information between…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Congqi Cao , Yue Lu , Yifan Zhang , Dongmei Jiang , Yanning Zhang

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

Recognizing human activities from visual inputs, particularly through a first-person viewpoint, is essential for enabling robots to replicate human behavior. Egocentric vision, characterized by cameras worn by observers, captures diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Amirshayan Nasirimajd , Chiara Plizzari , Simone Alberto Peirone , Marco Ciccone , Giuseppe Averta , Barbara Caputo

Attention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current research on video attention generally focuses on adopting a…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Yanbin Hao , Shuo Wang , Pei Cao , Xinjian Gao , Tong Xu , Jinmeng Wu , Xiangnan He

In this paper, we propose a new approach to under-stand actions in egocentric videos that exploits the semantics of object interactions at both frame and temporal levels. At the frame level, we use a region-based approach that takes as…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Alejandro Cartas , Petia Radeva , Mariella Dimiccoli