English
Related papers

Related papers: EgoIntrospect: An Egocentric Dataset and Benchmark…

200 papers

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Ziyang Wang , Yue Zhang , Shoubin Yu , Ce Zhang , Zengqi Zhao , Jaehong Yoon , Hyunji Lee , Gedas Bertasius , Mohit Bansal

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the…

Artificial Intelligence · Computer Science 2025-10-15 Nie Lin , Yansen Wang , Dongqi Han , Weibang Jiang , Jingyuan Li , Ryosuke Furuta , Yoichi Sato , Dongsheng Li

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yuan-Ming Li , Wei-Jin Huang , An-Lan Wang , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ahad Jawaid , Yu Xiang

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current methods are limited. Linear probes and other white-box methods compress high-dimensional…

Artificial Intelligence · Computer Science 2026-04-14 Nicolas Martorell , Bruno Bianchi

Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patterns. While synthetic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Rosario Leonardi , Francesco Ragusa , Daniele Materia , Alessandro Passanisi , James Fort , Jakob Engel , Giovanni Maria Farinella

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday…

First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of…

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid