中文
相关论文

相关论文: HCQA-1.5 @ Ego4D EgoSchema Challenge 2025

200 篇论文

In this technical report, we present our solution for the EgoPlan Challenge in ICML 2024. To address the real-world egocentric task planning problem, we introduce a novel planning framework which comprises three stages: long-term memory…

机器人学 · 计算机科学 2024-07-30 Letian Shi , Qi Lv , Xiang Deng , Liqiang Nie

This report presents the CuriosAI team's submission to the EgoExo4D Proficiency Estimation Challenge at CVPR 2025. We propose two methods for multi-view skill assessment: (1) a multi-task learning framework using Sapiens-2B that jointly…

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Sicheng Yang , Yukai Huang , Shitong Sun , Weitong Cai , Jiankang Deng , Jifei Song , Zhensong Zhang

Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the first…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ruiping Liu , Junwei Zheng , Yufan Chen , Di Wen , Shaofang Quan , Chengzhi Wu , Jiaming Zhang , Kailun Yang , Kunyu Peng , Rainer Stiefelhagen

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehensive understanding of video contents. Most previous works…

计算机视觉与模式识别 · 计算机科学 2022-08-11 Sipeng Zheng , Qi Zhang , Bei Liu , Qin Jin , Jianlong Fu

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

We introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets. We have developed a strong and robust question engine that leverages scene…

计算与语言 · 计算机科学 2019-07-12 Drew A. Hudson , Christopher D. Manning

We present EgoCOL, an egocentric camera pose estimation method for open-world 3D object localization. Our method leverages sparse camera pose reconstructions in a two-fold manner, video and scan independently, to estimate the camera pose of…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Cristhian Forigua , Maria Escobar , Jordi Pont-Tuset , Kevis-Kokitsi Maninis , Pablo Arbeláez

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Deheng Zhang , Yuqian Fu , Runyi Yang , Yang Miao , Tianwen Qian , Xu Zheng , Guolei Sun , Ajad Chhatkuli , Xuanjing Huang , Yu-Gang Jiang , Luc Van Gool , Danda Pani Paudel

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jinjie Mai , Abdullah Hamdi , Silvio Giancola , Chen Zhao , Bernard Ghanem

This report describes our submission to the Ego4D Moment Queries Challenge 2023. Our submission extends ActionFormer, a latest method for temporal action localization. Our extension combines an improved ground-truth assignment strategy…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Lin Sui , Fangzhou Mu , Yin Li

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

In this report, we present our approach for the Natural Language Query track and Goal Step track of the Ego4D Episodic Memory Benchmark at CVPR 2024. Both challenges require the localization of actions within long video sequences using…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Yisen Feng , Haoyu Zhang , Yuquan Xie , Zaijing Li , Meng Liu , Liqiang Nie

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Baoxiong Jia , Ting Lei , Song-Chun Zhu , Siyuan Huang

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

This paper presents a state-of-the-art model for visual question answering (VQA), which won the first place in the 2017 VQA Challenge. VQA is a task of significant importance for research in artificial intelligence, given its multimodal…

计算机视觉与模式识别 · 计算机科学 2017-08-10 Damien Teney , Peter Anderson , Xiaodong He , Anton van den Hengel