English
Related papers

Related papers: TouchAnything: A Dataset and Framework for Bimanua…

200 papers

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Vishakha Lall , Yisi Liu

Egocentric gesture recognition is a pivotal technology for enhancing natural human-computer interaction, yet traditional RGB-based solutions suffer from motion blur and illumination variations in dynamic scenarios. While event cameras show…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Luming Wang , Hao Shi , Xiaoting Yin , Kailun Yang , Kaiwei Wang , Jian Bai

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Karttikeya Mangalam , Raiymbek Akshulakov , Jitendra Malik

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

This work presents NeuroTouch, an optical-based tactile sensor that combines a highly deformable dome-shaped soft material with an integrated neuromorphic camera, leveraging frame-based and dynamic vision for gesture detection. Our approach…

Human-Computer Interaction · Computer Science 2025-02-03 Victor Hoffmann , Federico Paredes-Valles , Valentina Cavinato

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

On a minute-to-minute basis people undergo numerous fluid interactions with objects that barely register on a conscious level. Recent neuroscientific research demonstrates that humans have a fixed size prior for salient objects. This…

Computer Vision and Pattern Recognition · Computer Science 2015-11-10 Gedas Bertasius , Hyun Soo Park , Jianbo Shi

Equipping visitors of a cultural site with a wearable device allows to easily collect information about their preferences which can be exploited to improve the fruition of cultural goods with augmented reality. Moreover, egocentric video…

Computer Vision and Pattern Recognition · Computer Science 2020-02-04 Francesco Ragusa , Antonino Furnari , Sebastiano Battiato , Giovanni Signorello , Giovanni Maria Farinella

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

Multimedia · Computer Science 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what''…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Yuhang Yang , Wei Zhai , Chengfeng Wang , Chengjun Yu , Yang Cao , Zheng-Jun Zha

Action recognition is essential for egocentric video understanding, allowing automatic and continuous monitoring of Activities of Daily Living (ADLs) without user effort. Existing literature focuses on 3D hand pose input, which requires…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Wiktor Mucha , Martin Kampel

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

We present Ego-Only, the first approach that enables state-of-the-art action detection on egocentric (first-person) videos without any form of exocentric (third-person) transferring. Despite the content and appearance gap separating the two…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Huiyu Wang , Mitesh Kumar Singh , Lorenzo Torresani

Current upper extremity outcome measures for persons with cervical spinal cord injury (cSCI) lack the ability to directly collect quantitative information in home and community environments. A wearable first-person (egocentric) camera…

Human-Computer Interaction · Computer Science 2024-05-07 Jirapat Likitlersuang , Elizabeth R. Sumitro , Tianshi Cao , Ryan J. Visee , Sukhvinder Kalsi-Ryan , Jose Zariffa

Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high infrastructure costs, complex hardware dependencies, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Bowen Yang , Zishuo Li , Yang Sun , Changtao Miao , Yifan Yang , Man Luo , Xiaotong Yan , Feng Jiang , Jinchuan Shi , Yankai Fu , Ning Chen , Junkai Zhao , Pengwei Wang , Guocai Yao , Shanghang Zhang , Hao Chen , Zhe Li , Kai Zhu

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Brent Yi , Vickie Ye , Maya Zheng , Yunqi Li , Lea Müller , Georgios Pavlakos , Yi Ma , Jitendra Malik , Angjoo Kanazawa

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in egocentric settings and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Francesco Ragusa , Antonino Furnari , Salvatore Livatino , Giovanni Maria Farinella

Egocentric gestures are the most natural form of communication for humans to interact with wearable devices such as VR/AR helmets and glasses. A major issue in such scenarios for real-world applications is that may easily become necessary…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Zhengwei Wang , Qi She , Tejo Chalasani , Aljosa Smolic

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoCast, a bimodal method for 3D human pose forecasting using…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Maria Escobar , Juanita Puentes , Cristhian Forigua , Jordi Pont-Tuset , Kevis-Kokitsi Maninis , Pablo Arbelaez