English
Related papers

Related papers: Building Egocentric Procedural AI Assistant: Metho…

200 papers

A hallmark of advanced artificial intelligence is the capacity to progress from passive visual perception to the strategic modification of visual information to facilitate complex reasoning. This advanced capability, however, remains…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jingkun Ma , Runzhe Zhan , Yang Li , Di Sun , Hou Pong Chan , Lidia S. Chao , Derek F. Wong

Visual Planning for Assistance (VPA) aims to predict a sequence of user actions required to achieve a specified goal based on a video showing the user's progress. Although recent advances in multimodal large language models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Ce Zhang , Yale Song , Ruta Desai , Michael Louis Iuzzolino , Joseph Tighe , Gedas Bertasius , Satwik Kottur

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Shraman Pramanick , Yale Song , Sayan Nag , Kevin Qinghong Lin , Hardik Shah , Mike Zheng Shou , Rama Chellappa , Pengchuan Zhang

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Binzhu Xie , Shi Qiu , Sicheng Zhang , Yinqiao Wang , Hao Xu , Muzammal Naseer , Chi-Wing Fu , Pheng-Ann Heng

Vision-language models (VLMs) aligned with general human objectives, such as being harmless and hallucination-free, have become valuable assistants of humans in managing visual tasks. However, people with diversified backgrounds have…

Artificial Intelligence · Computer Science 2025-06-03 Yongqi Li , Shen Zhou , Xiaohu Li , Xin Miao , Jintao Wen , Mayi Xu , Jianhao Chen , Birong Pan , Hankun Kang , Yuanyuan Zhu , Ming Zhong , Tieyun Qian

This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled Task Guidance (PTG) program. This dataset comprises 3,355…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Brian VanVoorst , Nicholas Walczak , Christopher Gilleo , Charles Meissner , Fabio Felix , Iran Roman , Bea Steers , Claudio Silva , Yuhan Shen , Zijia Lu , Shih-Po Lee , Ehsan Elhamifar

Recent advances in the areas of multimodal machine learning and artificial intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision, Natural Language Processing, and Embodied AI. Whereas many…

Machine Learning · Computer Science 2022-05-26 Jonathan Francis , Nariaki Kitamura , Felix Labelle , Xiaopeng Lu , Ingrid Navarro , Jean Oh

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Egocentric human pose estimation aims to estimate human body poses and develop body representations from a first-person camera perspective. It has gained vast popularity in recent years because of its wide range of applications in sectors…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Md Mushfiqur Azam , Kevin Desai

The growing adoption of augmented and virtual reality (AR and VR) technologies in industrial training and on-the-job assistance has created new opportunities for intelligent, context-aware support systems. As workers perform complex tasks…

Human-Computer Interaction · Computer Science 2025-11-18 Mahya Qorbani , Kamran Paynabar , Mohsen Moghaddam

Natural interaction with virtual objects in AR/VR environments makes for a smooth user experience. Gestures are a natural extension from real world to augmented space to achieve these interactions. Finding discriminating spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Tejo Chalasani , Jan Ondrej , Aljosa Smolic

The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Yuqian Yuan , Ronghao Dang , Long Li , Wentong Li , Dian Jiao , Xin Li , Deli Zhao , Fan Wang , Wenqiao Zhang , Jun Xiao , Yueting Zhuang

Action recognition is essential for egocentric video understanding, allowing automatic and continuous monitoring of Activities of Daily Living (ADLs) without user effort. Existing literature focuses on 3D hand pose input, which requires…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Wiktor Mucha , Martin Kampel

Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patterns. While synthetic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Rosario Leonardi , Francesco Ragusa , Daniele Materia , Alessandro Passanisi , James Fort , Jakob Engel , Giovanni Maria Farinella

This paper addresses the daily challenges encountered by visually impaired individuals, such as limited access to information, navigation difficulties, and barriers to social interaction. To alleviate these challenges, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Inpyo Song , Minjun Joo , Joonhyung Kwon , Jangwon Lee

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower model and leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Baoqi Pei , Guo Chen , Jilan Xu , Yuping He , Yicheng Liu , Kanghua Pan , Yifei Huang , Yali Wang , Tong Lu , Limin Wang , Yu Qiao