中文
相关论文

相关论文: Rescaling Egocentric Vision

200 篇论文

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Data scarcity presents a key bottleneck for imitation learning in robotic manipulation. In this paper, we focus on random exploration data-actions and video sequences produced autonomously via motions to randomly sampled positions in the…

机器人学 · 计算机科学 2025-09-19 Shutong Jin , Axel Kaliff , Ruiyu Wang , Muhammad Zahid , Florian T. Pokorny

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach…

计算机视觉与模式识别 · 计算机科学 2023-01-06 Shuhan Tan , Tushar Nagarajan , Kristen Grauman

This report presents ContextRefine-CLIP (CR-CLIP), an efficient model for visual-textual multi-instance retrieval tasks. The approach is based on the dual-encoder AVION, on which we introduce a cross-modal attention flow module to achieve…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Jing He , Yiqing Wang , Lingling Li , Kexin Zhang , Puhua Chen

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

计算机视觉与模式识别 · 计算机科学 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multimodal large language models (MLLMs) are promising candidates,…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Haozhe Qi , Shaokai Ye , Alexander Mathis , Mackenzie W. Mathis

Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learning-based action…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Thanh-Dat Truong , Khoa Luu

Recently, attempts have been made to collect millions of videos to train CNN models for action recognition in videos. However, curating such large-scale video datasets requires immense human labor, and training CNNs on millions of videos…

计算机视觉与模式识别 · 计算机科学 2015-12-23 Shugao Ma , Sarah Adel Bargal , Jianming Zhang , Leonid Sigal , Stan Sclaroff

Counting in long videos remains a fundamental yet underexplored challenge in computer vision. Real-world recordings often span tens of minutes or longer and contain sparse, diverse events, making long-range temporal reasoning particularly…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fumihiko Tsuchiya , Taiki Miyanishi , Mahiro Ukai , Nakamasa Inoue , Shuhei Kurita , Yusuke Iwasawa , Yutaka Matsuo

Egocentric action anticipation aims to predict the future actions the camera wearer will perform from the observation of the past. While predictions about the future should be available before the predicted events take place, most…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Antonino Furnari , Giovanni Maria Farinella

We present the submission of Samsung AI Centre Cambridge to the CVPR2020 EPIC-Kitchens Action Recognition Challenge. In this challenge, action recognition is posed as the problem of simultaneously predicting a single `verb' and `noun' class…

计算机视觉与模式识别 · 计算机科学 2020-07-07 Juan-Manuel Perez-Rua , Antoine Toisoul , Brais Martinez , Victor Escorcia , Li Zhang , Xiatian Zhu , Tao Xiang

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action…

计算机视觉与模式识别 · 计算机科学 2021-01-20 Yu Wu , Linchao Zhu , Xiaohan Wang , Yi Yang , Fei Wu

In our recent dietary assessment field studies on passive dietary monitoring in Ghana, we have collected over 250k in-the-wild images. The dataset is an ongoing effort to facilitate accurate measurement of individual food and nutrient…

With the rapid increase of users of wearable cameras in recent years and of the amount of data they produce, there is a strong need for automatic retrieval and summarization techniques. This work addresses the problem of automatically…

计算机视觉与模式识别 · 计算机科学 2017-08-21 Aniol Lidon , Marc Bolaños , Mariella Dimiccoli , Petia Radeva , Maite Garolera , Xavier Giró-i-Nieto

Video Question Answering (VideoQA) is a task that requires a model to analyze and understand both the visual content given by the input video and the textual part given by the question, and the interaction between them in order to produce a…

计算机视觉与模式识别 · 计算机科学 2020-08-25 Alex Falcon , Oswald Lanz , Giuseppe Serra

Despite the recent progress on 6D object pose estimation methods for robotic grasping, a substantial performance gap persists between the capabilities of these methods on existing datasets and their efficacy in real-world grasping and…

机器人学 · 计算机科学 2024-12-18 Abdelrahman Younes , Tamim Asfour

Human activities exhibit a strong correlation between actions and the places where these are performed, such as washing something at a sink. More specifically, in daily living environments we may identify particular locations, hereinafter…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Simone Alberto Peirone , Gabriele Goletto , Mirco Planamente , Andrea Bottino , Barbara Caputo , Giuseppe Averta