中文
相关论文

相关论文: Do Egocentric Video-Language Models Truly Understa…

200 篇论文

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what''…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Yuhang Yang , Wei Zhai , Chengfeng Wang , Chengjun Yu , Yang Cao , Zheng-Jun Zha

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Binjie Zhang , Mike Zheng Shou

Pixel-level recognition of objects manipulated by the user from egocentric images enables key applications spanning assistive technologies, industrial safety, and activity monitoring. However, progress in this area is currently hindered by…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Nicola Messina , Rosario Leonardi , Luca Ciampi , Fabio Carrara , Giovanni Maria Farinella , Fabrizio Falchi , Antonino Furnari

We present a unified framework for understanding 3D hand and object interactions in raw image sequences from egocentric RGB cameras. Given a single RGB image, our model jointly estimates the 3D hand and object poses, models their…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Bugra Tekin , Federica Bogo , Marc Pollefeys

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective…

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited number of works have…

计算机视觉与模式识别 · 计算机科学 2019-10-16 Alejandro Cartas , Jordi Luque , Petia Radeva , Carlos Segura , Mariella Dimiccoli

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Dayou Li , Lulin Liu , Bangya Liu , Shijie Zhou , Jiu Feng , Ziqi Lu , Minghui Zheng , Chenyu You , Zhiwen Fan

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video,…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Jilan Xu , Yifei Huang , Baoqi Pei , Junlin Hou , Qingqiu Li , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Sheng Zhou , Junbin Xiao , Qingyun Li , Yicong Li , Xun Yang , Dan Guo , Meng Wang , Tat-Seng Chua , Angela Yao

Recent advances in multimodal large language models (MLLMs) offer a promising approach for natural language-based scene change queries in virtual reality (VR). Prior work on applying MLLMs for object state understanding has focused on…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Shiyi Ding , Shaoen Wu , Ying Chen

Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Tz-Ying Wu , Kyle Min , Subarna Tripathi , Nuno Vasconcelos

Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer. Current temporal action localization methods typically rely…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Erhang Zhang , Junyi Ma , Yin-Dong Zheng , Yixuan Zhou , Hesheng Wang

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Yisen Feng , Haoyu Zhang , Meng Liu , Weili Guan , Liqiang Nie

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Yifan Shen , Jiateng Liu , Xinzhuo Li , Yuanzhe Liu , Bingxuan Li , Houze Yang , Wenqi Jia , Yijiang Li , Tianjiao Yu , James Matthew Rehg , Xu Cao , Ismini Lourentzou

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scale and diversity, while using purely exocentric…

计算机视觉与模式识别 · 计算机科学 2021-04-19 Yanghao Li , Tushar Nagarajan , Bo Xiong , Kristen Grauman

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Vishakha Lall , Yisi Liu

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Ronghao Dang , Yuqian Yuan , Wenqi Zhang , Yifei Xin , Boqiang Zhang , Long Li , Liuyi Wang , Qinyang Zeng , Xin Li , Lidong Bing

Egocentric gesture recognition is a pivotal technology for enhancing natural human-computer interaction, yet traditional RGB-based solutions suffer from motion blur and illumination variations in dynamic scenarios. While event cameras show…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Luming Wang , Hao Shi , Xiaoting Yin , Kailun Yang , Kaiwei Wang , Jian Bai

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Wenqi Zhou , Kai Cao , Hao Zheng , Yunze Liu , Xinyi Zheng , Miao Liu , Per Ola Kristensson , Walterio Mayol-Cuevas , Fan Zhang , Weizhe Lin , Junxiao Shen

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata