English
Related papers

Related papers: EgoTV: Egocentric Task Verification from Natural L…

200 papers

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

Egocentric human video data, which captures rich human-environment interactions and can be collected at scale, has become a key driver of embodied intelligence research. However, existing egocentric datasets typically lack tactile sensing,…

Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is introduced to step-by-step support daily procedural tasks in a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Junlong Li , Huaiyuan Xu , Sijie Cheng , Kejun Wu , Kim-Hui Yap , Lap-Pui Chau , Yi Wang

Egocentric vision systems aim to understand the spatial surroundings and the wearer's behavior inside it, including motions, activities, and interactions. We argue that egocentric systems must additionally detect physiological states to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Björn Braun , Rayan Armani , Manuel Meier , Max Moebus , Christian Holz

Understanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person's emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Matthias Jammot , Björn Braun , Paul Streli , Rafael Wampfler , Christian Holz

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods, usually optimizing for high temporal…

Artificial Intelligence · Computer Science 2026-02-16 Zhaoyu Chen , Hongnan Lin , Yongwei Nie , Fei Ma , Xuemiao Xu , Fei Yu , Chengjiang Long

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Boshen Xu , Ziheng Wang , Yang Du , Zhinan Song , Sipeng Zheng , Qin Jin

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations - sparsity of action…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household tasks. However, VLM-based auto-labeling is often noisy because the primary data sources…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Lulin Liu , Dayou Li , Yiqing Liang , Sicong Jiang , Hitesh Vijay , Hezhen Hu , Xuhai Xu , Zirui Liu , Srinivas Shakkottai , Manling Li , Zhiwen Fan

Research on egocentric tasks in computer vision has mostly focused on head-mounted cameras, such as fisheye cameras or embedded cameras inside immersive headsets. We argue that the increasing miniaturization of optical sensors will lead to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Dominik Hollidt , Paul Streli , Jiaxi Jiang , Yasaman Haghighi , Changlin Qian , Xintong Liu , Christian Holz

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Ashwin Vinod , Shrey Pandit , Aditya Vavre , Linshen Liu

Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preferences, and routines, such personalization is essential for truly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Yanshuo Wang , Yuan Xu , Xuesong Li , Jie Hong , Yizhou Wang , Chang Wen Chen , Wentao Zhu

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Giuseppe Averta

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Gen Li , Kaifeng Zhao , Siwei Zhang , Xiaozhong Lyu , Mihai Dusmanu , Yan Zhang , Marc Pollefeys , Siyu Tang

We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view. While dense video captioning (predicting time segments and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Takehiko Ohkawa , Takuma Yagi , Taichi Nishimura , Ryosuke Furuta , Atsushi Hashimoto , Yoshitaka Ushiku , Yoichi Sato

The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentric videos and assign the corresponding verb--noun action label. In this report, we formulate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhiheng Fu , Zixu Li , Zhiwei Chen , Fangxu Liu , Yupeng Hu , Weili Guan , Liqiang Nie

In situated collaboration, speakers often use intentionally underspecified deictic commands (e.g., ``pass me \textit{that}''), whose referent becomes identifiable only by aligning speech with a brief co-speech pointing \emph{stroke}.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Weijie Zhou , Xuantang Xiong , Zhenlin Hu , Xiaomeng Zhu , Chaoyang Zhao , Honghui Dong , Zhengyou Zhang , Ming Tang , Jinqiao Wang

Vision-Language Models (VLMs) have shown great success as foundational models for downstream vision and natural language applications in a variety of domains. However, these models are limited to reasoning over objects and actions currently…

Robotics · Computer Science 2025-06-13 Zachary Chavis , Hyun Soo Park , Stephen J. Guy