English
Related papers

Related papers: Motion Focus Recognition in Fast-Moving Egocentric…

200 papers

Egocentric vision (a.k.a. first-person vision - FPV) applications have thrived over the past few years, thanks to the availability of affordable wearable cameras and large annotated datasets. The position of the wearable camera (usually…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Andrea Bandini , José Zariffa

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Egocentric vision holds great promises for increasing access to visual information and improving the quality of life for people with visual impairments, with object recognition being one of the daily challenges for this population. While we…

Computer Vision and Pattern Recognition · Computer Science 2020-03-02 Kyungjun Lee , Abhinav Shrivastava , Hernisa Kacorri

Advances in deep learning have enabled the development of models that have exhibited a remarkable tendency to recognize and even localize actions in videos. However, they tend to experience errors when faced with scenes or examples beyond…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Sathyanarayanan N. Aakur , Sanjoy Kundu , Nikhil Gunti

In the context of fitness coaching or for rehabilitation purposes, the motor actions of a human participant must be observed and analyzed for errors in order to provide effective feedback. This task is normally carried out by human coaches,…

Artificial Intelligence · Computer Science 2017-09-27 Felix Hülsmann , Stefan Kopp , Mario Botsch

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

With the widespread use of installed cameras, video-based monitoring approaches have seized considerable attention for different purposes like assisted living. Temporal redundancy and the sheer size of raw videos are the two most common…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Ali Abdari , Pouria Amirjan , Azadeh Mansouri

Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Mingfei Chen , Yifan Wang , Zhengqin Li , Homanga Bharadhwaj , Yujin Chen , Chuan Qin , Ziyi Kou , Yuan Tian , Eric Whitmire , Rajinder Sodhi , Hrvoje Benko , Eli Shlizerman , Yue Liu

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Evonne Ng , Donglai Xiang , Hanbyul Joo , Kristen Grauman

We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is crucial for augmented reality, human-computer interaction, assistive technologies, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Huy Anh Nguyen , Feras Dayoub , Minh Hoai

Action recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Xiao-Yu Zhang , Haichao Shi , Changsheng Li , Kai Zheng , Xiaobin Zhu , Lixin Duan

This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distinguish the relevant…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Hazel Doughty , Fida Mohammad Thoker , Cees G. M. Snoek

In this paper, we propose a novel self-supervised learning model for estimating continuous ego-motion from video. Our model learns to estimate camera motion by watching RGBD or RGB video streams and determining translational and rotation…

Computational Geometry · Computer Science 2018-06-28 Minhaeng Lee , Charless C. Fowlkes

While egocentric cameras like GoPro are gaining popularity, the videos they capture are long, boring, and difficult to watch from start to end. Fast forwarding (i.e. frame sampling) is a natural choice for faster video browsing. However,…

Computer Vision and Pattern Recognition · Computer Science 2017-01-04 Yair Poleg , Tavi Halperin , Chetan Arora , Shmuel Peleg

Learning an egocentric action recognition model from video data is challenging due to distractors (e.g., irrelevant objects) in the background. Further integrating object information into an action model is hence beneficial. Existing…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Victor Escorcia , Ricardo Guerrero , Xiatian Zhu , Brais Martinez

Planning at a higher level of abstraction instead of low level torques improves the sample efficiency in reinforcement learning, and computational efficiency in classical planning. We propose a method to learn such hierarchical…

Robotics · Computer Science 2019-10-16 Ashish Kumar , Saurabh Gupta , Jitendra Malik

Vision-language models (VLMs) have demonstrated excellent high-level planning capabilities, enabling locomotion skill learning from video demonstrations without the need for meticulous human-level reward design. However, the improper frame…

Robotics · Computer Science 2025-05-14 Xianghui Wang , Xinming Zhang , Yanjun Chen , Xiaoyu Shen , Wei Zhang

In this work, we address the problem of precisely localizing key frames of an action, for example, the precise time that a pitcher releases a baseball, or the precise time that a crowd begins to applaud. Key frame localization is a largely…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Iljung S. Kwak , Jian-Zhong Guo , Adam Hantman , David Kriegman , Kristin Branson

Many cameras implement auto-focus functionality. However, they typically require the user to manually identify the location to be focused on. While such an approach works for temporally-sparse autofocusing functionality (e.g., photo…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Wolfgang Fuhl , Thiago Santini , Enkelejda Kasneci