English
Related papers

Related papers: VidEgoThink: Assessing Egocentric Video Understand…

200 papers

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents…

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Ziyang Wang , Yue Zhang , Shoubin Yu , Ce Zhang , Zengqi Zhao , Jaehong Yoon , Hyunji Lee , Gedas Bertasius , Mohit Bansal

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jinzhao Li , Yinuo Chen , Dongxu Piao , Panwang Pan , Yifan Yu , Dong Wang , Honglei Yan , Liang Yue , Shaofei Wang , Yixin Chen , Siyuan Huang , Miao Liu

Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in Zero-Shot Egocentric Action Recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Guangzhao Dai , Xiangbo Shu , Wenhao Wu , Rui Yan , Jiachao Zhang

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Baoqi Pei , Yifei Huang , Jilan Xu , Yuping He , Guo Chen , Fei Wu , Yu Qiao , Jiangmiao Pang

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Rohit Sinha , Aditya Kanade , Sai Srinivas Kancheti , Vineeth N Balasubramanian , Tanuja Ganu

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Gen Luo , Ganlin Yang , Ziyang Gong , Guanzhou Chen , Haonan Duan , Erfei Cui , Ronglei Tong , Zhi Hou , Tianyi Zhang , Zhe Chen , Shenglong Ye , Lewei Lu , Jingbo Wang , Wenhai Wang , Jifeng Dai , Yu Qiao , Rongrong Ji , Xizhou Zhu

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Wenqi Zhou , Kai Cao , Hao Zheng , Yunze Liu , Xinyi Zheng , Miao Liu , Per Ola Kristensson , Walterio Mayol-Cuevas , Fan Zhang , Weizhe Lin , Junxiao Shen

Multimodal large language models (MLLMs) act as essential interfaces, connecting humans with AI technologies in multimodal applications. However, current MLLMs face challenges in accurately interpreting object orientation in images due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Ji Hyeok Jung , Eun Tae Kim , Seoyeon Kim , Joo Ho Lee , Bumsoo Kim , Buru Chang

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of…

The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Yi Chen , Yuying Ge , Yixiao Ge , Mingyu Ding , Bohao Li , Rui Wang , Ruifeng Xu , Ying Shan , Xihui Liu

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the two essential parts…

Artificial Intelligence · Computer Science 2023-06-13 Jinjie Mai , Jun Chen , Bing Li , Guocheng Qian , Mohamed Elhoseiny , Bernard Ghanem

Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (only video, or video-query pair where query is an image…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Raghav Goyal , Effrosyni Mavroudi , Xitong Yang , Sainbayar Sukhbaatar , Leonid Sigal , Matt Feiszli , Lorenzo Torresani , Du Tran

Multi-modal large language models (MLLMs) have achieved remarkable performance on objective multimodal perception tasks, but their ability to interpret subjective, emotionally nuanced multimodal content remains largely unexplored. Thus, it…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Qu Yang , Mang Ye , Bo Du

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chengjun Yu , Xuhan Zhu , Chaoqun Du , Pengfei Yu , Wei Zhai , Yang Cao , Zheng-Jun Zha

MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Tingyu Song , Tongyan Hu , Guo Gan , Yilun Zhao

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jianzhe Ma , Zhonghao Cao , Shangkui Chen , Yichen Xu , Wenxuan Wang , Qin Jin