English
Related papers

Related papers: EgoTextVQA: Towards Egocentric Scene-Text Aware Vi…

200 papers

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Deheng Zhang , Yuqian Fu , Runyi Yang , Yang Miao , Tianwen Qian , Xu Zheng , Guolei Sun , Ajad Chhatkuli , Xuanjing Huang , Yu-Gang Jiang , Luc Van Gool , Danda Pani Paudel

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jianzhe Ma , Zhonghao Cao , Shangkui Chen , Yichen Xu , Wenxuan Wang , Qin Jin

Video Question Answering (VideoQA) is a task that requires a model to analyze and understand both the visual content given by the input video and the textual part given by the question, and the interaction between them in order to produce a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Alex Falcon , Oswald Lanz , Giuseppe Serra

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Insu Lee , Wooje Park , Jaeyun Jang , Minyoung Noh , Kyuhong Shim , Byonghyo Shim

Existing efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decisionmaking and heavy reliance on scene-text recognition. In this paper, we propose to study Grounded TextVideoQA by forcing models to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Sheng Zhou , Junbin Xiao , Xun Yang , Peipei Song , Dan Guo , Angela Yao , Meng Wang , Tat-Seng Chua

As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Kangsan Kim , Yanlai Yang , Suji Kim , Woongyeong Yeo , Youngwan Lee , Mengye Ren , Sung Ju Hwang

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Karttikeya Mangalam , Raiymbek Akshulakov , Jitendra Malik

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ashish Seth , Xinhao Mei , Changsheng Zhao , Varun Nagaraja , Ernie Chang , Gregory P. Meyer , Gael Le Lan , Yunyang Xiong , Vikas Chandra , Yangyang Shi , Dinesh Manocha , Zhipeng Cai

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yanjun Li , Yuqian Fu , Tianwen Qian , Qi'ao Xu , Silong Dai , Danda Pani Paudel , Luc Van Gool , Xiaoling Wang

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Baoxiong Jia , Ting Lei , Song-Chun Zhu , Siyuan Huang

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

Machine Learning · Computer Science 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness. Scene text recognition in motion is a challenging problem, while textual cues typically appear for a short time span,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 George Tom , Minesh Mathew , Sergi Garcia , Dimosthenis Karatzas , C. V. Jawahar

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comprehension scheme for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Haoyu Zhang , Yuquan Xie , Yisen Feng , Zaijing Li , Meng Liu , Liqiang Nie

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

Artificial Intelligence · Computer Science 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu
‹ Prev 1 2 3 10 Next ›