English
Related papers

Related papers: Unlocking Exocentric Video-Language Data for Egoce…

200 papers

Object understanding in egocentric visual data is arguably a fundamental research topic in egocentric vision. However, existing object datasets are either non-egocentric or have limitations in object categories, visual content, and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chenchen Zhu , Fanyi Xiao , Andres Alvarado , Yasmine Babaei , Jiabo Hu , Hichem El-Mohri , Sean Chang Culatana , Roshan Sumbaly , Zhicheng Yan

Human emotion is expressed in many communication modalities and media formats and so their computational study is equally diversified into natural language processing, audio signal analysis, computer vision, etc. Similarly, the large…

Machine Learning · Computer Science 2023-08-16 Sven Buechel , Udo Hahn

Interactive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos. We demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-04-11 Mohit Goyal , Sahil Modi , Rishabh Goyal , Saurabh Gupta

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Humans can rearrange objects in cluttered environments using egocentric perception, navigating occlusions without global coordinates. Inspired by this capability, we study long-horizon multi-object non-prehensile rearrangement for mobile…

Robotics · Computer Science 2026-02-23 Boyuan An , Zhexiong Wang , Yipeng Wang , Jiaqi Li , Sihang Li , Jing Zhang , Chen Feng

Vision Language Models (VLMs) have achieved strong performance across diverse video understanding tasks. However, their viewpoint invariant training limits their ability to understand egocentric properties (e.g., human object interactions)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Dominick Reilly , Manish Kumar Govind , Le Xue , Srijan Das

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Hao Li , Qiao Sun

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Shraman Pramanick , Yale Song , Sayan Nag , Kevin Qinghong Lin , Hardik Shah , Mike Zheng Shou , Rama Chellappa , Pengchuan Zhang

We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Ronghao Dang , Yuqian Yuan , Yunxuan Mao , Kehan Li , Jiangpin Liu , Zhikai Wang , Xin Li , Fan Wang , Deli Zhao

From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Si-En Hong , James Tribble , Alexander Lake , Hao Wang , Chaoyi Zhou , Ashish Bastola , Siyu Huang , Eisa Chaudhary , Brian Canada , Ismahan Arslan-Ari , Abolfazl Razi

Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preferences, and routines, such personalization is essential for truly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Yanshuo Wang , Yuan Xu , Xuesong Li , Jie Hong , Yizhou Wang , Chang Wen Chen , Wentao Zhu

We introduce the Visual Experience Dataset (VEDB), a compilation of over 240 hours of egocentric video combined with gaze- and head-tracking data that offers an unprecedented view of the visual world as experienced by human observers. The…

Existing 3D facial emotion modeling have been constrained by limited emotion classes and insufficient datasets. This paper introduces "Emo3D", an extensive "Text-Image-Expression dataset" spanning a wide spectrum of human emotions, each…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Mahshid Dehghani , Amirahmad Shafiee , Ali Shafiei , Neda Fallah , Farahmand Alizadeh , Mohammad Mehdi Gholinejad , Hamid Behroozi , Jafar Habibi , Ehsaneddin Asgari

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Tushar Nagarajan , Kristen Grauman

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu

We study the understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Yixin Chen , Qing Li , Deqian Kong , Yik Lun Kei , Song-Chun Zhu , Tao Gao , Yixin Zhu , Siyuan Huang

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach…

Computer Vision and Pattern Recognition · Computer Science 2023-01-06 Shuhan Tan , Tushar Nagarajan , Kristen Grauman

Research on egocentric tasks in computer vision has mostly focused on head-mounted cameras, such as fisheye cameras or embedded cameras inside immersive headsets. We argue that the increasing miniaturization of optical sensors will lead to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Dominik Hollidt , Paul Streli , Jiaxi Jiang , Yasaman Haghighi , Changlin Qian , Xintong Liu , Christian Holz

Robot manipulation research still suffers from significant data scarcity: even the largest robot datasets are orders of magnitude smaller and less diverse than those that fueled recent breakthroughs in language and vision. We introduce…

Robotics · Computer Science 2026-05-29 Marion Lepert , Jiaying Fang , Jeannette Bohg

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yanjun Li , Yuqian Fu , Tianwen Qian , Qi'ao Xu , Silong Dai , Danda Pani Paudel , Luc Van Gool , Xiaoling Wang