English
Related papers

Related papers: Retrieval-Augmented Egocentric Video Captioning

200 papers

With the rapid development of wearable cameras, a massive collection of egocentric video for first-person visual perception becomes available. Using egocentric videos to predict first-person activity faces many challenges, including limited…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Linfeng Xu , Qingbo Wu , Lili Pan , Fanman Meng , Hongliang Li , Chiyuan He , Hanxin Wang , Shaoxu Cheng , Yu Dai

This paper presents a framework for recognition of human activity from egocentric video and eye tracking data obtained from a head-mounted eye tracker. Three channels of information such as eye movement, ego-motion, and visual features are…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Anjith George , Aurobinda Routray

In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause significant object…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Joohyun Chang , Soyeon Hong , Hyogun Lee , Seong Jong Ha , Dongho Lee , Seong Tae Kim , Jinwoo Choi

We pose keystep recognition as a node classification task, and propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos. Our approach,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge. Human-generated queries for video datasets `in the wild' vary a lot in terms of degree of specificity,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-17 Yang Liu , Samuel Albanie , Arsha Nagrani , Andrew Zisserman

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yanjun Li , Yuqian Fu , Tianwen Qian , Qi'ao Xu , Silong Dai , Danda Pani Paudel , Luc Van Gool , Xiaoling Wang

In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Wei Zhang , Bairui Wang , Lin Ma , Wei Liu

Pre-training on large scale unlabelled datasets has shown impressive performance improvements in the fields of computer vision and natural language processing. Given the advent of large-scale instructional video datasets, a common strategy…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Valentin Gabeur , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Vishakha Lall , Yisi Liu

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

Robotics · Computer Science 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

Computer Vision and Pattern Recognition · Computer Science 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

Multimodal large language models (MLLMs) act as essential interfaces, connecting humans with AI technologies in multimodal applications. However, current MLLMs face challenges in accurately interpreting object orientation in images due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Ji Hyeok Jung , Eun Tae Kim , Seoyeon Kim , Joo Ho Lee , Bumsoo Kim , Buru Chang

Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as…

Robotics · Computer Science 2025-11-10 Yichen Zhu , Feifei Feng

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Artificial Intelligence · Computer Science 2025-06-09 Yichi Zhang , Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , Seungwhan Moon

Wearable cameras offer a hands-free way to record egocentric images of daily experiences, where social events are of special interest. The first step towards detection of social events is to track the appearance of multiple persons involved…

Computer Vision and Pattern Recognition · Computer Science 2017-01-24 Maedeh Aghaei , Mariella Dimiccoli , Petia Radeva

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Shraman Pramanick , Yale Song , Sayan Nag , Kevin Qinghong Lin , Hardik Shah , Mike Zheng Shou , Rama Chellappa , Pengchuan Zhang

We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various applications, however, it remains largely under-explored. It has been…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Taeho Kang , Kyungjin Lee , Jinrui Zhang , Youngki Lee

Can Video-LLMs achieve consistent temporal understanding when videos capture the same event from different viewpoints? To study this, we introduce EgoExo-Con (Consistency), a benchmark of comprehensively synchronized egocentric and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Minjoon Jung , Junbin Xiao , Junghyun Kim , Byoung-Tak Zhang , Angela Yao

Understanding users' activities from head-mounted cameras is a fundamental task for Augmented and Virtual Reality (AR/VR) applications. A typical approach is to train a classifier in a supervised manner using data labeled by humans. This…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Satoshi Tsutsui , Ruta Desai , Karl Ridgeway