English
Related papers

Related papers: Look and Tell: A Dataset for Multimodal Grounding …

200 papers

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video synthesis and dynamic scene understanding. The dataset contains nearly 1,000 short egocentric videos captured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jae Yong Lee , Daniel Scharstein , Akash Bapat , Hao Hu , Andrew Fu , Haoru Zhao , Paul Sammut , Xiang Li , Stephen Jeapes , Anik Gupta , Lior David , Saketh Madhuvarasu , Jay Girish Joshi , Jason Wither

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

We introduce a method for egocentric videoconferencing that enables hands-free video calls, for instance by people wearing smart glasses or other mixed-reality devices. Videoconferencing portrays valuable non-verbal communication and face…

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integration with natural…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in…

We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Eadom Dessalene , Botao He , Michael Maynord , Yonatan Tussa , Pavan Mantripragada , Yianni Karabati , Nirupam Roy , Yiannis Aloimonos

The ability to associate touch with sight is essential for tasks that require physically interacting with objects in the world. We propose a dataset with paired visual and tactile data called Touch and Go, in which human data collectors…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Fengyu Yang , Chenyang Ma , Jiacheng Zhang , Jing Zhu , Wenzhen Yuan , Andrew Owens

We study the understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Yixin Chen , Qing Li , Deqian Kong , Yik Lun Kei , Song-Chun Zhu , Tao Gao , Yixin Zhu , Siyuan Huang

Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Mingfei Chen , Yifan Wang , Zhengqin Li , Homanga Bharadhwaj , Yujin Chen , Chuan Qin , Ziyi Kou , Yuan Tian , Eric Whitmire , Rajinder Sodhi , Hrvoje Benko , Eli Shlizerman , Yue Liu

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Ling Li , Bowen Liu , Zinuo Zhan , Peng Jie , Jianhui Zhong , Kenglun Chang , Zhidong Deng

To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to…

Spatial reasoning poses a particular challenge for intelligent agents and is at the same time a prerequisite for their successful interaction and communication in the physical world. One such reasoning task is to describe the position of a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Kyra Ahrens , Matthias Kerzel , Jae Hee Lee , Cornelius Weber , Stefan Wermter

The widespread adoption of commercial autonomous vehicles (AVs) and advanced driver assistance systems (ADAS) may largely depend on their acceptance by society, for which their perceived trustworthiness and interpretability to riders are…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Enna Sachdeva , Nakul Agarwal , Suhas Chundi , Sean Roelofs , Jiachen Li , Mykel Kochenderfer , Chiho Choi , Behzad Dariush

This paper addresses the task of assessing PICU team's leadership skills by developing an automated analysis framework based on egocentric vision. We identify key behavioral cues, including fixation object, eye contact, and conversation…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Liangyang Ouyang , Yuki Sakai , Ryosuke Furuta , Hisataka Nozawa , Hikoro Matsui , Yoichi Sato

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

Understanding transcripts of immersive multimodal conversations is challenging because speakers frequently rely on visual context and non-verbal cues, such as gestures and visual attention, which are not captured in speech alone. This lack…

Human-Computer Interaction · Computer Science 2025-09-11 Riccardo Bovo , Frederik Brudy , George Fitzmaurice , Fraser Anderson

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Advanced in-cabin sensing technologies, especially vision based approaches, have tremendously progressed user interaction inside the vehicle, paving the way for new applications of natural user interaction. Just as humans use multiple modes…

Human-Computer Interaction · Computer Science 2022-02-16 Abdul Rafey Aftab , Michael von der Beeck