English
Related papers

Related papers: Fine-grained Spatiotemporal Grounding on Egocentri…

200 papers

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Yisen Feng , Haoyu Zhang , Meng Liu , Weili Guan , Liqiang Nie

Visual object tracking is a key component to many egocentric vision problems. However, the full spectrum of challenges of egocentric tracking faced by an embodied AI is underrepresented in many existing datasets; these tend to focus on…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Hao Tang , Kevin Liang , Matt Feiszli , Weiyao Wang

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (only video, or video-query pair where query is an image…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Raghav Goyal , Effrosyni Mavroudi , Xitong Yang , Sainbayar Sukhbaatar , Leonid Sigal , Matt Feiszli , Lorenzo Torresani , Du Tran

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Zihui Xue , Kristen Grauman

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Brian Chen , Nina Shvetsova , Andrew Rouditchenko , Daniel Kondermann , Samuel Thomas , Shih-Fu Chang , Rogerio Feris , James Glass , Hilde Kuehne

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Taewoong Kang , Kinam Kim , Dohyeon Kim , Minho Park , Junha Hyung , Jaegul Choo

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Baoqi Pei , Yifei Huang , Jilan Xu , Yuping He , Guo Chen , Fei Wu , Yu Qiao , Jiangmiao Pang

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Xiaoqi Wang , Yi Wang , Lap-Pui Chau

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Qi'ao Xu , Tianwen Qian , Yuqian Fu , Kailing Li , Yang Jiao , Jiacheng Zhang , Xiaoling Wang , Liang He

Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring to light that current egocentric video question-answering…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Chiara Plizzari , Alessio Tonioni , Yongqin Xian , Achin Kulshrestha , Federico Tombari

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentation and tracking in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Yash Bhalgat , Vadim Tschernezki , Iro Laina , João F. Henriques , Andrea Vedaldi , Andrew Zisserman

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Baoqi Pei , Yifei Huang , Jilan Xu , Guo Chen , Yuping He , Lijin Yang , Yali Wang , Weidi Xie , Yu Qiao , Fei Wu , Limin Wang

Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patterns. While synthetic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Rosario Leonardi , Francesco Ragusa , Daniele Materia , Alessandro Passanisi , James Fort , Jakob Engel , Giovanni Maria Farinella

Video grounding aims to localize a spatio-temporal section in a video corresponding to an input text query. This paper addresses a critical limitation in current video grounding methodologies by introducing an Open-Vocabulary…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Syed Talal Wasim , Muzammal Naseer , Salman Khan , Ming-Hsuan Yang , Fahad Shahbaz Khan
‹ Prev 1 2 3 10 Next ›