English
Related papers

Related papers: EgoVideo: Exploring Egocentric Foundation Model an…

200 papers

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Guo Chen , Sen Xing , Zhe Chen , Yi Wang , Kunchang Li , Yizhuo Li , Yi Liu , Jiahao Wang , Yin-Dong Zheng , Bingkun Huang , Zhiyu Zhao , Junting Pan , Yifei Huang , Zun Wang , Jiashuo Yu , Yinan He , Hongjie Zhang , Tong Lu , Yali Wang , Limin Wang , Yu Qiao

This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC VQA challenge. HD-EPIC evaluates whether a vision-language model can reason over realistic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhiwei Chen , Yupeng Hu , Zixu Li , Zhiheng Fu , Guozhi Qiu , Weili Guan , Liqiang Nie

This technical report describes the EgoTask Translation approach that explores relations among a set of egocentric video tasks in the Ego4D challenge. To improve the primary task of interest, we propose to leverage existing models developed…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Tz-Ying Wu , Kyle Min , Subarna Tripathi , Nuno Vasconcelos

In this report, we present our champion solution for Ego4D Natural Language Queries (NLQ) Challenge in CVPR 2023. Essentially, to accurately ground in a video, an effective egocentric feature extractor and a powerful grounding model are…

Computer Vision and Pattern Recognition · Computer Science 2023-06-28 Zhijian Hou , Lei Ji , Difei Gao , Wanjun Zhong , Kun Yan , Chao Li , Wing-Kwong Chan , Chong-Wah Ngo , Nan Duan , Mike Zheng Shou

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Kevin Qinghong Lin , Alex Jinpeng Wang , Rui Yan , Eric Zhongcong Xu , Rongcheng Tu , Yanru Zhu , Wenzhe Zhao , Weijie Kong , Chengfei Cai , Hongfa Wang , Wei Liu , Mike Zheng Shou

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Yisen Feng , Haoyu Zhang , Qiaohui Chu , Meng Liu , Weili Guan , Yaowei Wang , Liqiang Nie

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehensive understanding of video contents. Most previous works…

Computer Vision and Pattern Recognition · Computer Science 2022-08-11 Sipeng Zheng , Qi Zhang , Bei Liu , Qin Jin , Jianlong Fu

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Weikun Peng , Denys Iliash , Manolis Savva

The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentric videos and assign the corresponding verb--noun action label. In this report, we formulate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhiheng Fu , Zixu Li , Zhiwei Chen , Fangxu Liu , Yupeng Hu , Weili Guan , Liqiang Nie

We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS evaluation benchmark, we include 9x more points that go…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Ahmad Darkhalil , Rhodri Guerrier , Adam W. Harley , Dima Damen

To enable progress towards egocentric agents capable of understanding everyday tasks specified in natural language, we propose a benchmark and a synthetic dataset called Egocentric Task Verification (EgoTV). The goal in EgoTV is to verify…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Rishi Hazra , Brian Chen , Akshara Rai , Nitin Kamra , Ruta Desai

In this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023. The aim of this task is to predict a sequence of future actions that will take place at an arbitrary time or…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Tatsuya Ishibashi , Kosuke Ono , Noriyuki Kugo , Yuji Sato
‹ Prev 1 2 3 10 Next ›