English
Related papers

Related papers: Masked Autoencoders for Egocentric Video Understan…

200 papers

This report describes our submission called "TarHeels" for the Ego4D: Object State Change Classification Challenge. We use a transformer-based video recognition model and leverage the Divided Space-Time Attention mechanism for classifying…

Computer Vision and Pattern Recognition · Computer Science 2023-01-05 Md Mohaiminul Islam , Gedas Bertasius

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

This technical report describes the EgoTask Translation approach that explores relations among a set of egocentric video tasks in the Ego4D challenge. To improve the primary task of interest, we propose to leverage existing models developed…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

In this report, we present the transferring pretrained video mask autoencoders(VideoMAE) to egocentric tasks for Ego4d Looking at me Challenge. VideoMAE is the data-efficient pretraining model for self-supervised video pre-training and can…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Yinan He , Guo Chen

We implemented Video Swin Transformer as a base architecture for the tasks of Point-of-No-Return temporal localization and Object State Change Classification. Our method achieved competitive performance on both challenges.

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Maria Escobar , Laura Daza , Cristina González , Jordi Pont-Tuset , Pablo Arbeláez

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Yisen Feng , Haoyu Zhang , Qiaohui Chu , Meng Liu , Weili Guan , Yaowei Wang , Liqiang Nie

Capturing the state changes of interacting objects is a key technology for understanding human-object interactions. This technical report describes our method using heterogeneous backbones for the Ego4D Object State Change Classification…

Computer Vision and Pattern Recognition · Computer Science 2022-11-17 Yin-Dong Zheng , Guo Chen , Jiahao Wang , Tong Lu , Limin Wang

Generic Event Boundary Detection (GEBD) tasks aim at detecting generic, taxonomy-free event boundaries that segment a whole video into chunks. In this paper, we apply Masked Autoencoders to improve algorithm performance on the GEBD tasks.…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Rui He , Yuanxi Sun , Youzeng Li , Zuwei Huang , Feng Hu , Xu Cheng , Jie Tang

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-channel) audio…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Guo Chen , Sen Xing , Zhe Chen , Yi Wang , Kunchang Li , Yizhuo Li , Yi Liu , Jiahao Wang , Yin-Dong Zheng , Bingkun Huang , Zhiyu Zhao , Junting Pan , Yifei Huang , Zun Wang , Jiashuo Yu , Yinan He , Hongjie Zhang , Tong Lu , Yali Wang , Limin Wang , Yu Qiao

We study the problem of unsupervised domain adaptation for egocentric videos. We propose a transformer-based model to learn class-discriminative and domain-invariant feature representations. It consists of two novel designs. The first…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Xiaoyu Zhu , Junwei Liang , Po-Yao Huang , Alex Hauptmann

Masked autoencoding has become a successful pretraining paradigm for Transformer models for text, images, and, recently, point clouds. Raw automotive datasets are suitable candidates for self-supervised pre-training as they generally are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Georg Hess , Johan Jaxing , Elias Svensson , David Hagerman , Christoffer Petersson , Lennart Svensson

This report describes our submission to the Ego4D Moment Queries Challenge 2023. Our submission extends ActionFormer, a latest method for temporal action localization. Our extension combines an improved ground-truth assignment strategy…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Lin Sui , Fangzhou Mu , Yin Li

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

The recently released Ego4D dataset and benchmark significantly scales and diversifies the first-person visual perception data. In Ego4D, the Visual Queries 2D Localization task aims to retrieve objects appeared in the past from the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Mengmeng Xu , Cheng-Yang Fu , Yanghao Li , Bernard Ghanem , Juan-Manuel Perez-Rua , Tao Xiang

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze serves as a key non-verbal communication cue that reflects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Wei-Cheng Lin , Chih-Ming Lien , Chen Lo , Chia-Hung Yeh
‹ Prev 1 2 3 10 Next ›