English
Related papers

Related papers: EgoLife: Towards Egocentric Life Assistant

200 papers

Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yogesh Kulkarni , Pooyan Fazli

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Qiaohui Chu , Haoyu Zhang , Meng Liu , Yisen Feng , Haoxiang Shi , Liqiang Nie

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Giuseppe Averta

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

Surgical tool detection is a fundamental task for understanding egocentric open surgery videos. However, detecting surgical tools presents significant challenges due to their highly imbalanced class distribution, similar shapes and similar…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Ryo Fujii , Hideo Saito , Hiroki Kajita

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. LifelongMemory generates concise video activity descriptions of the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Ying Wang , Yanlai Yang , Mengye Ren

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs often lack the social awareness to discern when to intervene as…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xijun Wang , Tanay Sharma , Achin Kulshrestha , Abhimitra Meka , Aveek Purohit , Dinesh Manocha

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the mainstream.…

Computer Vision and Pattern Recognition · Computer Science 2015-01-14 Sven Bambach

Both assistant driving and self-driving have attracted a great amount of attention in the last few years. However, the majority of research efforts focus on safe driving; few research has been conducted on in-vehicle climate control, or…

Machine Learning · Computer Science 2020-06-17 Feng Hu

Improving human health and well-being requires an accurate and effective understanding of an individual's physical and mental state throughout daily life. To support this goal, we utilized smartphones, smartwatches, and sleep sensors to…

Signal Processing · Electrical Eng. & Systems 2025-08-07 Se Won Oh , Hyuntae Jeong , Seungeun Chung , Jeong Mook Lim , Kyoung Ju Noh , Sunkyung Lee , Gyuwon Jung

We present a video summarization approach for egocentric or "wearable" camera data. Given hours of video, the proposed method produces a compact storyboard summary of the camera wearer's day. In contrast to traditional keyframe selection…

Computer Vision and Pattern Recognition · Computer Science 2015-05-20 Yong Jae Lee , Kristen Grauman

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Video Question Answering (VideoQA) is a task that requires a model to analyze and understand both the visual content given by the input video and the textual part given by the question, and the interaction between them in order to produce a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Alex Falcon , Oswald Lanz , Giuseppe Serra

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric…

Human-Computer Interaction · Computer Science 2026-04-10 Valdemar Danry , Javier Hernandez , Andrew Wilson , Pattie Maes , Judith Amores

We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on…

Psychological counseling is a highly personalized and dynamic process that requires therapists to continuously monitor emotional changes, document session insights, and maintain therapeutic continuity. In this paper, we introduce…

Human-Computer Interaction · Computer Science 2025-04-24 Xianghe Liu , Jiaqi Xu , Tao Sun

The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wild datasets align…

Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Chentao Li , Zirui Gao , Mingze Gao , Yinglian Ren , Jianjiang Feng , Jie Zhou
‹ Prev 1 8 9 10 Next ›