English
Related papers

Related papers: PhysBrain: Human Egocentric Data as a Bridge from …

200 papers

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jaeyoung Choi , Hyeondong Kim , Yujin Kim , Daehee Park

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Gen Li , Kaifeng Zhao , Siwei Zhang , Xiaozhong Lyu , Mihai Dusmanu , Yan Zhang , Marc Pollefeys , Siyu Tang

Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work, we introduce EmbodiedGPT, an end-to-end multi-modal…

Robotics · Computer Science 2023-09-15 Yao Mu , Qinglong Zhang , Mengkang Hu , Wenhai Wang , Mingyu Ding , Jun Jin , Bin Wang , Jifeng Dai , Yu Qiao , Ping Luo

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Artificial Intelligence · Computer Science 2025-05-21 Joel Currie , Gioele Migno , Enrico Piacenti , Maria Elena Giannaccini , Patric Bach , Davide De Tommaso , Agnieszka Wykowska

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

Human children far exceed modern machine learning algorithms in their sample efficiency, achieving high performance in key domains with much less data than current models. This ''data gap'' is a key challenge both for building intelligent…

Deploying humanoid robots in real-world settings is fundamentally challenging, as it demands tight integration of perception, locomotion, and manipulation under partial-information observations and dynamically changing environments. As well…

Robotics · Computer Science 2026-02-05 Yu Bai , MingMing Yu , Chaojie Li , Ziyi Bai , Xinlong Wang , Börje F. Karlsson

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation,…

Robotics · Computer Science 2026-04-07 Yifu Yuan , Haiqin Cui , Yaoting Huang , Yibin Chen , Fei Ni , Zibin Dong , Pengyi Li , Yan Zheng , Hongyao Tang , Jianye Hao

Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs…

Human-Computer Interaction · Computer Science 2025-05-26 Arnav Verma , Kushin Mukherjee , Christopher Potts , Elisa Kreiss , Judith E. Fan

Generating instructional images of human daily actions from an egocentric viewpoint serves as a key step towards efficient skill transfer. In this paper, we introduce a novel problem -- egocentric action frame generation. The goal is to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Bolin Lai , Xiaoliang Dai , Lawrence Chen , Guan Pang , James M. Rehg , Miao Liu

Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric perspective, which serves as a key bridge between humans and machines in action…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mengmeng Ge , Takashi Isobe , Xu Jia , Yanan Sun , Zetong Yang , Weinong Wang , Dong Zhou , Dong Li , Huchuan Lu , Emad Barsoum

Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Ruibing Hou , Mingyue Zhou , Yuwei Gui , Mingshuang Luo , Bingpeng Ma , Hong Chang , Shiguang Shan , Xilin Chen

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Heeseung Yun , Joonil Na , Jaeyeon Kim , Calvin Murdock , Gunhee Kim

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Haoyu Zhen , Xiaowen Qiu , Peihao Chen , Jincheng Yang , Xin Yan , Yilun Du , Yining Hong , Chuang Gan

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Wei Chow , Jiageng Mao , Boyi Li , Daniel Seita , Vitor Guizilini , Yue Wang

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Mohsen Gholami , Ahmad Rezaei , Zhou Weimin , Sitong Mao , Shunbo Zhou , Yong Zhang , Mohammad Akbari