English
Related papers

Related papers: Do You See What I Am Pointing At? Gesture-Based Eg…

200 papers

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Chentao Li , Zirui Gao , Mingze Gao , Yinglian Ren , Jianjiang Feng , Jie Zhou

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ye Pan , Chi Kit Wong , Yuanhuiyi Lyu , Hanqian Li , Jiahao Huo , Jiacheng Chen , Lutao Jiang , Xu Zheng , Xuming Hu

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Sheng Zhou , Junbin Xiao , Qingyun Li , Yicong Li , Xun Yang , Dan Guo , Meng Wang , Tat-Seng Chua , Angela Yao

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Ling Li , Bowen Liu , Zinuo Zhan , Peng Jie , Jianhui Zhong , Kenglun Chang , Zhidong Deng

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Baoxiong Jia , Ting Lei , Song-Chun Zhu , Siyuan Huang

Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Boshen Xu , Ziheng Wang , Yang Du , Zhinan Song , Sipeng Zheng , Qin Jin

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jaeyoung Choi , Hyeondong Kim , Yujin Kim , Daehee Park

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jianzhe Ma , Zhonghao Cao , Shangkui Chen , Yichen Xu , Wenxuan Wang , Qin Jin

Video Question Answering (VideoQA) is a task that requires a model to analyze and understand both the visual content given by the input video and the textual part given by the question, and the interaction between them in order to produce a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Alex Falcon , Oswald Lanz , Giuseppe Serra

Natural interaction with virtual objects in AR/VR environments makes for a smooth user experience. Gestures are a natural extension from real world to augmented space to achieve these interactions. Finding discriminating spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Tejo Chalasani , Jan Ondrej , Aljosa Smolic

Mobile virtual reality (VR) head mounted displays (HMD) have become popular among consumers in recent years. In this work, we demonstrate real-time egocentric hand gesture detection and localization on mobile HMDs. Our main contributions…

Computer Vision and Pattern Recognition · Computer Science 2017-12-15 Rohit Pandey , Marie White , Pavel Pidlypenskyi , Xue Wang , Christine Kaeser-Chen

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Qiaohui Chu , Haoyu Zhang , Meng Liu , Yisen Feng , Haoxiang Shi , Liqiang Nie

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang
‹ Prev 1 2 3 10 Next ›