English
Related papers

Related papers: The Toybox Dataset of Egocentric Visual Object Tra…

200 papers

Tolerance to image variations (e.g. translation, scale, pose, illumination) is an important desired property of any object recognition system, be it human or machine. Moving towards increasingly bigger datasets has been trending in computer…

Computer Vision and Pattern Recognition · Computer Science 2016-01-27 Ali Borji , Saeed Izadi , Laurent Itti

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

Machine learning and computer vision methods have a major impact on the study of natural animal behavior, as they enable the (semi-)automatic analysis of vast amounts of video data. Mice are the standard mammalian model system in most…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Patrik Reiske , Marcus N. Boon , Niek Andresen , Sole Traverso , Katharina Hohlbaum , Lars Lewejohann , Christa Thöne-Reineke , Olaf Hellwich , Henning Sprekeler

Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs) are essential to…

Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person egocentric vision…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Matteo Dunnhofer , Zaira Manigrasso , Christian Micheloni

We report on an extensive study of the benefits and limitations of current deep learning approaches to object recognition in robot vision scenarios, introducing a novel dataset used for our investigation. To avoid the biases in currently…

Robotics · Computer Science 2021-08-25 Giulia Pasquale , Carlo Ciliberto , Francesca Odone , Lorenzo Rosasco , Lorenzo Natale

This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled Task Guidance (PTG) program. This dataset comprises 3,355…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Brian VanVoorst , Nicholas Walczak , Christopher Gilleo , Charles Meissner , Fabio Felix , Iran Roman , Bea Steers , Claudio Silva , Yuhan Shen , Zijia Lu , Shih-Po Lee , Ehsan Elhamifar

Understanding how images of objects and scenes behave in response to specific ego-motions is a crucial aspect of proper visual development, yet existing visual learning methods are conspicuously disconnected from the physical source of…

Computer Vision and Pattern Recognition · Computer Science 2016-03-30 Dinesh Jayaraman , Kristen Grauman

Wearable cameras capture a first-person view of the daily activities of the camera wearer, offering a visual diary of the user behaviour. Detection of the appearance of people the camera user interacts with for social interactions analysis…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Estefania Talavera , Alexandre Cola , Nicolai Petkov , Petia Radeva

We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Eadom Dessalene , Botao He , Michael Maynord , Yonatan Tussa , Pavan Mantripragada , Yianni Karabati , Nirupam Roy , Yiannis Aloimonos

We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view. While dense video captioning (predicting time segments and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Takehiko Ohkawa , Takuma Yagi , Taichi Nishimura , Ryosuke Furuta , Atsushi Hashimoto , Yoshitaka Ushiku , Yoichi Sato

With the increasing performance of machine learning techniques in the last few years, the computer vision and robotics communities have created a large number of datasets for benchmarking object recognition tasks. These datasets cover a…

Computer Vision and Pattern Recognition · Computer Science 2016-11-18 Philipp Jund , Nichola Abdo , Andreas Eitel , Wolfram Burgard

Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Haisheng Su , Feixiang Song , Cong Ma , Wei Wu , Junchi Yan

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models train on raw, uniform visual inputs collected from…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Timothy Schaumlöffel , Arthur Aubret , Gemma Roig , Jochen Triesch

We propose a method to learn image representations from uncurated videos. We combine a supervised loss from off-the-shelf object detectors and self-supervised losses which naturally arise from the video-shot-frame-object hierarchy present…

Computer Vision and Pattern Recognition · Computer Science 2021-02-10 Rob Romijnders , Aravindh Mahendran , Michael Tschannen , Josip Djolonga , Marvin Ritter , Neil Houlsby , Mario Lucic

Image classification models built into visual support systems and other assistive devices need to provide accurate predictions about their environment. We focus on an application of assistive technology for people with visual impairments,…

Computer Vision and Pattern Recognition · Computer Science 2019-01-04 Marcus Klasson , Cheng Zhang , Hedvig Kjellström

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Rawal Khirodkar , Aayush Bansal , Lingni Ma , Richard Newcombe , Minh Vo , Kris Kitani

The binding problem in human cognition, concerning how the brain represents and connects objects within a fixed network of neural connections, remains a subject of intense debate. Most machine learning efforts addressing this issue in an…

Machine Learning · Computer Science 2023-10-18 Sindy Löwe , Phillip Lippe , Francesco Locatello , Max Welling