English
Related papers

Related papers: Rescaling Egocentric Vision

200 papers

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Data scarcity presents a key bottleneck for imitation learning in robotic manipulation. In this paper, we focus on random exploration data-actions and video sequences produced autonomously via motions to randomly sampled positions in the…

Robotics · Computer Science 2025-09-19 Shutong Jin , Axel Kaliff , Ruiyu Wang , Muhammad Zahid , Florian T. Pokorny

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach…

Computer Vision and Pattern Recognition · Computer Science 2023-01-06 Shuhan Tan , Tushar Nagarajan , Kristen Grauman

This report presents ContextRefine-CLIP (CR-CLIP), an efficient model for visual-textual multi-instance retrieval tasks. The approach is based on the dual-encoder AVION, on which we introduce a cross-modal attention flow module to achieve…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Jing He , Yiqing Wang , Lingling Li , Kexin Zhang , Puhua Chen

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multimodal large language models (MLLMs) are promising candidates,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Haozhe Qi , Shaokai Ye , Alexander Mathis , Mackenzie W. Mathis

Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learning-based action…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Thanh-Dat Truong , Khoa Luu

Recently, attempts have been made to collect millions of videos to train CNN models for action recognition in videos. However, curating such large-scale video datasets requires immense human labor, and training CNNs on millions of videos…

Computer Vision and Pattern Recognition · Computer Science 2015-12-23 Shugao Ma , Sarah Adel Bargal , Jianming Zhang , Leonid Sigal , Stan Sclaroff

Counting in long videos remains a fundamental yet underexplored challenge in computer vision. Real-world recordings often span tens of minutes or longer and contain sparse, diverse events, making long-range temporal reasoning particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fumihiko Tsuchiya , Taiki Miyanishi , Mahiro Ukai , Nakamasa Inoue , Shuhei Kurita , Yusuke Iwasawa , Yutaka Matsuo

Egocentric action anticipation aims to predict the future actions the camera wearer will perform from the observation of the past. While predictions about the future should be available before the predicted events take place, most…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Antonino Furnari , Giovanni Maria Farinella

We present the submission of Samsung AI Centre Cambridge to the CVPR2020 EPIC-Kitchens Action Recognition Challenge. In this challenge, action recognition is posed as the problem of simultaneously predicting a single `verb' and `noun' class…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Juan-Manuel Perez-Rua , Antoine Toisoul , Brais Martinez , Victor Escorcia , Li Zhang , Xiatian Zhu , Tao Xiang

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Kevin Qinghong Lin , Alex Jinpeng Wang , Rui Yan , Eric Zhongcong Xu , Rongcheng Tu , Yanru Zhu , Wenzhe Zhao , Weijie Kong , Chengfei Cai , Hongfa Wang , Wei Liu , Mike Zheng Shou

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action…

Computer Vision and Pattern Recognition · Computer Science 2021-01-20 Yu Wu , Linchao Zhu , Xiaohan Wang , Yi Yang , Fei Wu

In our recent dietary assessment field studies on passive dietary monitoring in Ghana, we have collected over 250k in-the-wild images. The dataset is an ongoing effort to facilitate accurate measurement of individual food and nutrient…

With the rapid increase of users of wearable cameras in recent years and of the amount of data they produce, there is a strong need for automatic retrieval and summarization techniques. This work addresses the problem of automatically…

Computer Vision and Pattern Recognition · Computer Science 2017-08-21 Aniol Lidon , Marc Bolaños , Mariella Dimiccoli , Petia Radeva , Maite Garolera , Xavier Giró-i-Nieto

Video Question Answering (VideoQA) is a task that requires a model to analyze and understand both the visual content given by the input video and the textual part given by the question, and the interaction between them in order to produce a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Alex Falcon , Oswald Lanz , Giuseppe Serra

Despite the recent progress on 6D object pose estimation methods for robotic grasping, a substantial performance gap persists between the capabilities of these methods on existing datasets and their efficacy in real-world grasping and…

Robotics · Computer Science 2024-12-18 Abdelrahman Younes , Tamim Asfour

Human activities exhibit a strong correlation between actions and the places where these are performed, such as washing something at a sink. More specifically, in daily living environments we may identify particular locations, hereinafter…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Simone Alberto Peirone , Gabriele Goletto , Mirco Planamente , Andrea Bottino , Barbara Caputo , Giuseppe Averta
‹ Prev 1 4 5 6 7 8 10 Next ›