English
Related papers

Related papers: Ego-EXTRA: video-language Egocentric Dataset for E…

200 papers

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang

We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video synthesis and dynamic scene understanding. The dataset contains nearly 1,000 short egocentric videos captured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jae Yong Lee , Daniel Scharstein , Akash Bapat , Hao Hu , Andrew Fu , Haoru Zhao , Paul Sammut , Xiang Li , Stephen Jeapes , Anik Gupta , Lior David , Saketh Madhuvarasu , Jay Girish Joshi , Jason Wither

Object understanding in egocentric visual data is arguably a fundamental research topic in egocentric vision. However, existing object datasets are either non-egocentric or have limitations in object categories, visual content, and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chenchen Zhu , Fanyi Xiao , Andres Alvarado , Yasmine Babaei , Jiabo Hu , Hichem El-Mohri , Sean Chang Culatana , Roshan Sumbaly , Zhicheng Yan

Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Ege Özsoy , Arda Mamur , Felix Tristram , Chantal Pellegrini , Magdalena Wysocki , Benjamin Busam , Nassir Navab

This paper addresses the daily challenges encountered by visually impaired individuals, such as limited access to information, navigation difficulties, and barriers to social interaction. To alleviate these challenges, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Inpyo Song , Minjun Joo , Joonhyung Kwon , Jangwon Lee

Accurately forecasting human trajectories from an egocentric perspective plays a central role in applications such as humanoid robotics, wearable sensing systems, and assistive navigation. However, progress in this direction remains limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Ahmad Yehia , Abduallah Mohamed , Tianyi Wang , Jiseop Byeon , Kun Qian , Junfeng Jiao , Christian Claudel

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury

Understanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person's emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Matthias Jammot , Björn Braun , Paul Streli , Rafael Wampfler , Christian Holz

Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in…

Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing benchmarks typically evaluate these abilities in isolation,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Jiaqi Yan , Ruilong Ren , Jingren Liu , Shuning Xu , Ling Wang , Yiheng Wang , Xinlin Zhong , Yun Wang , Long Zhang , Xiangyu Chen , Changzhi Sun , Jixiang Luo , Dell Zhang , Hao Sun , Chi Zhang , Xuelong Li

Mistake action detection is crucial for developing intelligent archives that detect workers' errors and provide feedback. Existing studies have focused on visually apparent mistakes in free-style activities, resulting in video-only…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Yuto Haneji , Taichi Nishimura , Hirotaka Kameko , Keisuke Shirai , Tomoya Yoshida , Keiya Kajimura , Koki Yamamoto , Taiyu Cui , Tomohiro Nishimoto , Shinsuke Mori

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Egocentric video is increasingly used as a data source for robot learning, activity understanding, and embodied AI research, but collecting it at scale remains fragmented in practice: each candidate host device, such as an Android phone,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Liuchuan Yu , Erdem Murat , Beichen Wang , Yan Zeng , Tingting Luo , Huizhen Zhou , Shanghao Li , Huining Feng , Zhigen Zhao , Ning Yang , Ke Jing , Yunhao Liu , Ruoya Sheng

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Kevin Qinghong Lin , Alex Jinpeng Wang , Rui Yan , Eric Zhongcong Xu , Rongcheng Tu , Yanru Zhu , Wenzhe Zhao , Weijie Kong , Chengfei Cai , Hongfa Wang , Wei Liu , Mike Zheng Shou

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

Machine Learning · Computer Science 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang