中文
相关论文

相关论文: HCQA-1.5 @ Ego4D EgoSchema Challenge 2025

200 篇论文

Visual question answering (VQA) systems are emerging from a desire to empower users to ask any natural language question about visual content and receive a valid answer in response. However, close examination of the VQA problem reveals an…

人工智能 · 计算机科学 2016-08-30 Danna Gurari , Kristen Grauman

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

In this report, we present the ReLER@ZJU1 submission to the Ego4D Moment Queries Challenge in ECCV 2022. In this task, the goal is to retrieve and localize all instances of possible activities in egocentric videos. Ego4D dataset is…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Jiayi Shao , Xiaohan Wang , Yi Yang

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences.…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Yuping He , Yifei Huang , Guo Chen , Baoqi Pei , Jilan Xu , Tong Lu , Jiangmiao Pang

We present MCQA, a learning-based algorithm for multimodal question answering. MCQA explicitly fuses and aligns the multimodal input (i.e. text, audio, and video), which forms the context for the query (question and answer). Our approach…

计算与语言 · 计算机科学 2020-04-28 Abhishek Kumar , Trisha Mittal , Dinesh Manocha

Embodied Question Answering (EQA) is a recently proposed task, where an agent is placed in a rich 3D environment and must act based solely on its egocentric input to answer a given question. The desired outcome is that the agent learns to…

计算机视觉与模式识别 · 计算机科学 2019-08-15 Cătălina Cangea , Eugene Belilovsky , Pietro Liò , Aaron Courville

This report presents our team's PCIE_Interaction solution for the Ego4D Social Interaction Challenge at CVPR 2025, addressing both Looking At Me (LAM) and Talking To Me (TTM) tasks. The challenge requires accurate detection of social…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Kanokphan Lertniphonphan , Feng Chen , Junda Xu , Fengbu Lan , Jun Xie , Tao Zhang , Zhepeng Wang

We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric video timestamp, the task requires anticipating the next…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Qiaohui Chu , Haoyu Zhang , Yisen Feng , Meng Liu , Weili Guan , Dongmei Jiang , Liqiang Nie

Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, existing MCQA benchmarks…

计算与语言 · 计算机科学 2025-06-02 Olga Loginova , Oleksandr Bezrukov , Ravi Shekhar , Alexey Kravets

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scale and diversity, while using purely exocentric…

计算机视觉与模式识别 · 计算机科学 2021-04-19 Yanghao Li , Tushar Nagarajan , Bo Xiong , Kristen Grauman

We propose a novel and challenging benchmark, AutoEval-Video, to comprehensively evaluate large vision-language models in open-ended video question answering. The comprehensiveness of AutoEval-Video is demonstrated in two aspects: 1)…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Xiuyuan Chen , Yuan Lin , Yuchen Zhang , Weiran Huang

With the development of eXtended Reality (XR), photo capturing and display technology based on head-mounted displays (HMDs) have experienced significant advancements and gained considerable attention. Egocentric spatial images and videos…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Xilei Zhu , Liu Yang , Huiyu Duan , Xiongkuo Min , Guangtao Zhai , Patrick Le Callet

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Binjie Zhang , Mike Zheng Shou

The HLTCOE Evaluation team participated in TREC VQA's Answer Generation (AG) task, for which we developed a listwise learning framework that aims to improve semantic precision and ranking consistency in answer generation. Given a…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Dengjia Zhang , Charles Weng , Katherine Guerrerio , Yi Lu , Kenton Murray , Alexander Martin , Reno Kriz , Benjamin Van Durme

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoCast, a bimodal method for 3D human pose forecasting using…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Maria Escobar , Juanita Puentes , Cristhian Forigua , Jordi Pont-Tuset , Kevis-Kokitsi Maninis , Pablo Arbelaez

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their ability to reason…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Shitong Sun , Ke Han , Yukai Huang , Weitong Cai , Jifei Song

In this chapter, we describe our question answering system, which was the winning system at the Human-Computer Question Answering (HCQA) Competition at the Thirty-first Annual Conference on Neural Information Processing Systems (NIPS). The…

计算与语言 · 计算机科学 2018-03-26 Ikuya Yamada , Ryuji Tamaki , Hiroyuki Shindo , Yoshiyasu Takefuji

Deep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of…

多媒体 · 计算机科学 2023-03-08 Weikai Kong , Shuhong Ye , Chenglin Yao , Jianfeng Ren