English
Related papers

Related papers: GPT4Ego: Unleashing the Potential of Pre-trained M…

200 papers

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Ashwin Vinod , Shrey Pandit , Aditya Vavre , Linshen Liu

Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Sijie Cheng , Zhicheng Guo , Jingwen Wu , Kechen Fang , Peng Li , Huaping Liu , Yang Liu

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Predicting pedestrian behavior is the key to ensure safety and reliability of autonomous vehicles. While deep learning methods have been promising by learning from annotated video frame sequences, they often fail to fully grasp the dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Jia Huang , Peng Jiang , Alvika Gautam , Srikanth Saripalli

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang

Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs) trained on curated…

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Mohsen Gholami , Ahmad Rezaei , Zhou Weimin , Sitong Mao , Shunbo Zhou , Yong Zhang , Mohammad Akbari

AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Alessandro Suglia , Claudio Greco , Katie Baker , Jose L. Part , Ioannis Papaioannou , Arash Eshghi , Ioannis Konstas , Oliver Lemon

Vision-Language Models (VLMs) have shown great success as foundational models for downstream vision and natural language applications in a variety of domains. However, these models are limited to reasoning over objects and actions currently…

Robotics · Computer Science 2025-06-13 Zachary Chavis , Hyun Soo Park , Stephen J. Guy

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in…

Eye gaze, encompassing fixations and saccades, provides critical insights into human intentions and future actions. This study introduces a gaze-regularized framework that enhances Vision Language Models (VLMs) for egocentric behavior…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Anupam Pani , Yanchao Yang

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot…

Sound · Computer Science 2024-11-20 Satvik Dixit , Laurie M. Heller , Chris Donahue

Vision-language models (VLMs) have demonstrated remarkable performance across various visual tasks, leveraging joint learning of visual and textual representations. While these models excel in zero-shot image tasks, their application to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Massimo Bosetti , Shibingfeng Zhang , Benedetta Liberatori , Giacomo Zara , Elisa Ricci , Paolo Rota

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang
‹ Prev 1 2 3 10 Next ›