English
Related papers

Related papers: Deep Dual Relation Modeling for Egocentric Interac…

200 papers

We present ongoing work to harness biological approaches to achieving highly efficient egocentric perception by combining the space-variant imaging architecture of the mammalian retina with Deep Learning methods. By pre-processing images…

Computer Vision and Pattern Recognition · Computer Science 2018-09-06 Nina Hristozova , Piotr Ozimek , Jan Paul Siebert

We consider the problem of localizing visitors in a cultural site from egocentric (first person) images. Localization information can be useful both to assist the user during his visit (e.g., by suggesting where to go and what to see next)…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Francesco Ragusa , Antonino Furnari , Sebastiano Battiato , Giovanni Signorello , Giovanni Maria Farinella

First-person interaction recognition is a challenging task because of unstable video conditions resulting from the camera wearer's movement. For human interaction recognition from a first-person viewpoint, this paper proposes a three-stream…

Computer Vision and Pattern Recognition · Computer Science 2020-02-20 Ye-Ji Kim , Dong-Gyu Lee , Seong-Whan Lee

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Wiktor Mucha , Florin Cuconasu , Naome A. Etori , Valia Kalokyri , Giovanni Trappolini

Action anticipation in egocentric videos is a difficult task due to the inherently multi-modal nature of human actions. Additionally, some actions happen faster or slower than others depending on the actor or surrounding context which could…

Computer Vision and Pattern Recognition · Computer Science 2021-09-03 Nada Osman , Guglielmo Camporese , Pasquale Coscia , Lamberto Ballan

Video action recognition, a critical problem in video understanding, has been gaining increasing attention. To identify actions induced by complex object-object interactions, we need to consider not only spatial relations among objects in a…

Computer Vision and Pattern Recognition · Computer Science 2019-05-08 Hao Huang , Luowei Zhou , Wei Zhang , Jason J. Corso , Chenliang Xu

Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capability requires instance…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Yunhan Zhao , Haoyu Ma , Shu Kong , Charless Fowlkes

In this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhaofeng Shi , Heqian Qiu , Lanxiao Wang , Fanman Meng , Qingbo Wu , Hongliang Li

The ability of robots to model their own dynamics is key to autonomous planning and learning, as well as for autonomous damage detection and recovery. Traditionally, dynamic models are pre-programmed or learned from external observations.…

Robotics · Computer Science 2024-03-19 Yuhang Hu , Boyuan Chen , Hod Lipson

Wearable egocentric cameras and machine learning have the potential to provide clinicians with a more nuanced understanding of patient hand use at home after stroke and spinal cord injury (SCI). However, they require detailed contextual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Adesh Kadambi , José Zariffa

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and…

Multimedia · Computer Science 2016-01-26 Sibo Song , Ngai-Man Cheung , Vijay Chandrasekhar , Bappaditya Mandal , Jie Lin

Large Language Models (LLMs) have made progress in various real-world tasks, which stimulates requirements for the evaluation of LLMs. Existing LLM evaluation methods are mainly supervised signal-based which depends on static datasets and…

Computation and Language · Computer Science 2023-09-11 Jiatong Li , Rui Li , Qi Liu

Learning an egocentric action recognition model from video data is challenging due to distractors (e.g., irrelevant objects) in the background. Further integrating object information into an action model is hence beneficial. Existing…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Victor Escorcia , Ricardo Guerrero , Xiatian Zhu , Brais Martinez

Egocentric vision holds great promises for increasing access to visual information and improving the quality of life for people with visual impairments, with object recognition being one of the daily challenges for this population. While we…

Computer Vision and Pattern Recognition · Computer Science 2020-03-02 Kyungjun Lee , Abhinav Shrivastava , Hernisa Kacorri

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Zhengyi Luo , Ryo Hachiuma , Ye Yuan , Kris Kitani

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

Multimedia · Computer Science 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

Egocentric vision aims to capture and analyse the world from the first-person perspective. We explore the possibilities for egocentric wearable devices to improve and enhance industrial use cases w.r.t. data collection, annotation,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Vivek Chavan , Oliver Heimann , Jörg Krüger

Accurately identifying hands in images is a key sub-task for human activity understanding with wearable first-person point-of-view cameras. Traditional hand segmentation approaches rely on a large corpus of manually labeled data to generate…

Computer Vision and Pattern Recognition · Computer Science 2018-06-18 Yubo Zhang , Vishnu Naresh Boddeti , Kris M. Kitani