English
Related papers

Related papers: MobileEgo Anywhere: Open Infrastructure for long h…

200 papers

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Yifei Huang , Guo Chen , Jilan Xu , Mingfang Zhang , Lijin Yang , Baoqi Pei , Hongjie Zhang , Lu Dong , Yali Wang , Limin Wang , Yu Qiao

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Liang Xu , Chengqun Yang , Zili Lin , Fei Xu , Yifan Liu , Congsheng Xu , Yiyi Zhang , Jie Qin , Xingdong Sheng , Yunhui Liu , Xin Jin , Yichao Yan , Wenjun Zeng , Xiaokang Yang

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not only increases…

Robotics · Computer Science 2026-04-10 Yanwen Zou , Chenyang Shi , Wenye Yu , Han Xue , Jun Lv , Ye Pan , Chuan Wen , Cewu Lu

Human children far exceed modern machine learning algorithms in their sample efficiency, achieving high performance in key domains with much less data than current models. This ''data gap'' is a key challenge both for building intelligent…

We address the challenge of predicting human visual attention during real-world navigation by measuring and modeling egocentric pedestrian eye gaze in an outdoor campus setting. We introduce the EgoCampus dataset, which spans 25 unique…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Ronan John , Aditya Kesari , Vincenzo DiMatteo , Kristin Dana

Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, have the potential…

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ashish Seth , Xinhao Mei , Changsheng Zhao , Varun Nagaraja , Ernie Chang , Gregory P. Meyer , Gael Le Lan , Yunyang Xiong , Vikas Chandra , Yangyang Shi , Dinesh Manocha , Zhipeng Cai

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially grounded driving actions by leveraging multimodal inputs,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Xingcheng Zhou , Xuyuan Han , Feng Yang , Yunpu Ma , Volker Tresp , Alois Knoll

Data-driven navigation algorithms are critically dependent on large-scale, high-quality real-world data collection for successful training and robust performance in realistic and uncontrolled conditions. To enhance the growing family of…

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data…

Robotics · Computer Science 2026-03-24 Xinhai Sun , Xiang Shi , Menglin Zou , Wenlong Huang

Robust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently…

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-person viewpoint changes that only move the camera,…

The 1st Cross-Domain EgoCross Challenge at EgoVis, CVPR 2026 evaluates whether multimodal large language models can reason over egocentric videos across surgery, industry, extreme sports, and animal perspective. We achieved second place in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Zixu Li , Zhiwei Chen , Zhiheng Fu , Wenbo Wang , Yupeng Hu , Weili Guan , Liqiang Nie

What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overload, smart glasses coupled with AI agents could turn the web…

Human-Computer Interaction · Computer Science 2026-03-03 Sicheng Yang , Yukai Huang , Weitong Cai , Shitong Sun , Fengyi Fang , You He , Yiqiao Xie , Jiankang Deng , Hang Zhang , Jifei Song , Zhensong Zhang

Imitation learning for manipulation has a well-known data scarcity problem. Unlike natural language and 2D computer vision, there is no Internet-scale corpus of data for dexterous manipulation. One appealing option is egocentric human…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ryan Hoque , Peide Huang , David J. Yoon , Mouli Sivapurapu , Jian Zhang

Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures primarily rely on linear or flat storage, lacking structural…

Robotics · Computer Science 2026-05-13 Yanbin Hu , Jin Cui , Jiayi Lu , Ruixuan Yang , Jun Ye , Boran Zhao , Xingyu Chen , Xuguang Lan , Pengju Ren

We present a novel system for real-time tracking of facial expressions using egocentric views captured from a set of infrared cameras embedded in a virtual reality (VR) headset. Our technology facilitates any user to accurately drive the…

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera…