Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we introduce EgoIntrospect, the first egocentric dataset captured in user-driven scenarios with self-annotations that explicitly reveal users' interactive intentions with AI assistants. EgoIntrospect was collected using a cross-device setup, providing synchronized video, audio, gaze, motion, and physiological signals. It consists of 180 hours of recordings from 60 subjects, with an average recording duration of 3 hours per subject. Leveraging EgoIntrospect, we formalize a suite of tasks centered on user internal states, including affective experience, interactive intent, and cognitive memory. We further process the annotations to construct benchmarks that evaluate the ability of modern multimodal large language models to reason about users' internal states from egocentric observations. Experiments on our benchmark suggest that existing multimodal large language models struggle to effectively leverage multimodal signals to infer users' subjective internal states. The dataset and annotations will be made publicly available to advance research in egocentric vision and wearable AI assistants. Project page: https://ego-introspect.github.io/
@article{arxiv.2605.17262,
title = {EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning},
author = {Zeyu Wang and Chang Liu and Eduardus Tjitrahardja and Yuntao Wang and Borislav Pavlov and Fangfei Gou and Jose Manuel Davila and Dai Shi and Ran Xu and Yue Pan and Jiayi Tan and Shuting Chang and Qi Wang and Jinzhao Li and Jiacheng Hua and Yifei Huang and Jingwei Sun and Yu Zhang and Liuxin Zhang and Guocai Yao and Jia Jia and Yin Li and Qianying Wang and Yuanchun Shi and Miao Liu},
journal= {arXiv preprint arXiv:2605.17262},
year = {2026}
}