English
Related papers

Related papers: EgoThinker: Unveiling Egocentric Reasoning with Sp…

200 papers

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-language models (LVLMs) often address this challenge by externalizing reasoning through…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yiming Liang , Yixiao Chen , Yiyang Zhou , Yixuan Wang , Shoubin Yu , Andong Deng , Fuxiao Liu , Qin Zhang , Chen Chen , Mohit Bansal , Huaxiu Yao

Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static scene assumption in…

Computer Vision and Pattern Recognition · Computer Science 2019-10-04 Jia-Wang Bian , Zhichao Li , Naiyan Wang , Huangying Zhan , Chunhua Shen , Ming-Ming Cheng , Ian Reid

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Marc Bolaños , Álvaro Peris , Francisco Casacuberta , Sergi Soler , Petia Radeva

Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social robots and human-computer interaction, where understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 He Hu , Tengjin Weng , Zebang Cheng , Yu Wang , Jiachen Luo , Björn Schuller , Zheng Lian , Laizhong Cui

In this study, we investigate various computer vision paradigms - supervised learning, unsupervised learning, and prompt fine-tuning - by assessing their ability to understand and interpret egocentric video data. Specifically, we examine…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Daniel Wen

When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion between third-person and first-person perspectives and a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Baiyu Pan , Daqin Luo , Junpeng Yang , Jiyuan Wang , Yixuan Zhang , Hailin Shi , Jichao Jiao

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Insu Lee , Wooje Park , Jaeyun Jang , Minyoung Noh , Kyuhong Shim , Byonghyo Shim

Multimodal large language models (MLLMs) act as essential interfaces, connecting humans with AI technologies in multimodal applications. However, current MLLMs face challenges in accurately interpreting object orientation in images due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Ji Hyeok Jung , Eun Tae Kim , Seoyeon Kim , Joo Ho Lee , Bumsoo Kim , Buru Chang

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Kun Ouyang , Yuanxin Liu , Haoning Wu , Yi Liu , Hao Zhou , Jie Zhou , Fandong Meng , Xu Sun

Smart glass is emerging as an useful device since it provides plenty of insights under hands-busy, eyes-on-task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Taegyoon Yoon , Yegyu Han , Seojin Ji , Jaewoo Park , Sojeong Kim , Taein Kwon , Hyung-Sin Kim

Large language models (LLMs) have shown impressive capabilities, but still struggle with complex reasoning tasks requiring multiple steps. While prompt-based methods like Chain-of-Thought (CoT) can improve LLM reasoning at inference time,…

Artificial Intelligence · Computer Science 2024-11-25 Haolin Chen , Yihao Feng , Zuxin Liu , Weiran Yao , Akshara Prabhakar , Shelby Heinecke , Ricky Ho , Phil Mui , Silvio Savarese , Caiming Xiong , Huan Wang

In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. LifelongMemory generates concise video activity descriptions of the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Ying Wang , Yanlai Yang , Mengye Ren

Capturing interaction of hands with objects is important to autonomously detect human actions from egocentric videos. In this work, we present a pyramid video transformer with a dynamic class token generator for egocentric action…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Chenbin Pan , Zhiqi Zhang , Senem Velipasalar , Yi Xu

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, especially when guided by explicit chain-of-thought (CoT) reasoning that verbalizes intermediate steps. While CoT improves both interpretability and accuracy,…

Multimodal large language models (LLMs) have made rapid progress in visual understanding, yet their extension from images to videos often reduces to a naive concatenation of frame tokens. In this work, we investigate what video finetuning…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ruiqi Yang , Tian Yun , Zihan Wang , Ellie Pavlick

The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Kejian Zhu , Zhuoran Jin , Hongbang Yuan , Jiachun Li , Shangqing Tu , Pengfei Cao , Yubo Chen , Kang Liu , Jun Zhao

Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in Zero-Shot Egocentric Action Recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Guangzhao Dai , Xiangbo Shu , Wenhao Wu , Rui Yan , Jiachao Zhang

Temporal Video Grounding (TVG), which requires pinpointing relevant temporal segments from video based on language query, has always been a highly challenging task in the field of video understanding. Videos often have a larger volume of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Feng Yue , Zhaoxing Zhang , Junming Jiao , Zhengyu Liang , Shiwen Cao , Feifei Zhang , Rong Shen

Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Mingfei Chen , Yifan Wang , Zhengqin Li , Homanga Bharadhwaj , Yujin Chen , Chuan Qin , Ziyi Kou , Yuan Tian , Eric Whitmire , Rajinder Sodhi , Hrvoje Benko , Eli Shlizerman , Yue Liu
‹ Prev 1 8 9 10 Next ›