中文
相关论文

相关论文: Do Egocentric Video-Language Models Truly Understa…

200 篇论文

Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Quankai Gao , Jiawei Yang , Qiangeng Xu , Le Chen , Yue Wang

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs often lack the social awareness to discern when to intervene as…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xijun Wang , Tanay Sharma , Achin Kulshrestha , Abhimitra Meka , Aveek Purohit , Dinesh Manocha

Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yogesh Kulkarni , Pooyan Fazli

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans'…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Danny Tran , Roberto Martín-Martín , Kristen Grauman

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2)…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Xiaoqi Wang , Yi Wang , Lap-Pui Chau

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Liang Xu , Chengqun Yang , Zili Lin , Fei Xu , Yifan Liu , Congsheng Xu , Yiyi Zhang , Jie Qin , Xingdong Sheng , Yunhui Liu , Xin Jin , Yichao Yan , Wenjun Zeng , Xiaokang Yang

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

Video diffusion models have recently achieved remarkable progress in realism and controllability. However, achieving seamless video translation across different perspectives, such as first-person (egocentric) and third-person (exocentric),…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Quanjian Song , Yiren Song , Kelly Peng , Yuan Gao , Mike Zheng Shou

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower model and leveraging…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Baoqi Pei , Guo Chen , Jilan Xu , Yuping He , Yicheng Liu , Kanghua Pan , Yifei Huang , Yali Wang , Tong Lu , Limin Wang , Yu Qiao

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is introduced to step-by-step support daily procedural tasks in a…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Junlong Li , Huaiyuan Xu , Sijie Cheng , Kejun Wu , Kim-Hui Yap , Lap-Pui Chau , Yi Wang

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transformer module to reason about each acting hand, and its…

计算机视觉与模式识别 · 计算机科学 2022-01-14 Jian Ma , Dima Damen

Recently, there has been a growing interest in analyzing human daily activities from data collected by wearable cameras. Since the hands are involved in a vast set of daily tasks, detecting hands in egocentric images is an important step…

计算机视觉与模式识别 · 计算机科学 2017-09-11 Alejandro Cartas , Mariella Dimiccoli , Petia Radeva

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

We present Ego-Only, the first approach that enables state-of-the-art action detection on egocentric (first-person) videos without any form of exocentric (third-person) transferring. Despite the content and appearance gap separating the two…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Huiyu Wang , Mitesh Kumar Singh , Lorenzo Torresani

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained…

计算机视觉与模式识别 · 计算机科学 2022-05-19 Alex Jinpeng Wang , Yixiao Ge , Guanyu Cai , Rui Yan , Xudong Lin , Ying Shan , Xiaohu Qie , Mike Zheng Shou