English
Related papers

Related papers: Spherical World-Locking for Audio-Visual Localizat…

200 papers

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video…

Artificial Intelligence · Computer Science 2026-05-14 Qinchuan Cheng , Zhantao Gong , Pengzhan Sun , Angela Yao , Xulei Yang , Shijie Li

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jingyu Li , Junjie Wu , Dongnan Hu , Xiangkai Huang , Bin Sun , Zhihui Hao , Xianpeng Lang , Xiatian Zhu , Li Zhang

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaffold rather than an explicit mapping objective. Starting from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xinyi Zheng , Yunze Liu , Chi-Hao Wu , Fan Zhang , Hao Zheng , Wenqi Zhou , Walterio W. Mayol-Cuevas , Junxiao Shen

Human and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Xinyu Yi , Yuxiao Zhou , Marc Habermann , Vladislav Golyanik , Shaohua Pan , Christian Theobalt , Feng Xu

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Weizhen Wang , Chenda Duan , Zhenghao Peng , Yuxin Liu , Bolei Zhou

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Insu Lee , Wooje Park , Jaeyun Jang , Minyoung Noh , Kyuhong Shim , Byonghyo Shim

Vision-language models (VLMs) demonstrate strong image-level scene understanding but often lack persistent memory, explicit spatial representations, and computational efficiency when reasoning over long video sequences. We present VL-KnG, a…

We propose Wolf, a WOrLd summarization Framework for accurate video captioning. Wolf is an automated captioning framework that adopts a mixture-of-experts approach, leveraging complementary strengths of Vision Language Models (VLMs). By…

We built our pipeline EgoLoc-v1, mainly inspired by EgoLoc. We propose a model ensemble strategy to improve the camera pose estimation part of the VQ3D task, which has been proven to be essential in previous work. The core idea is not only…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jinjie Mai , Abdullah Hamdi , Silvio Giancola , Chen Zhao , Bernard Ghanem

Panoramic distortion poses a significant challenge in 360 depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Junsong Zhang , Zisong Chen , Chunyu Lin , Lang Nie , Zhijie Shen , Kang Liao , Junda Huang , Yao Zhao

We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jack Hong , Shilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie

The ability to understand long videos is vital for embodied intelligent agents, because their effectiveness depends on how well they can accumulate, organize, and leverage long-horizon perceptual memories. Recently, multimodal LLMs have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Tatiana Zemskova , Solomon Andryushenko , Ilya Obrubov , Viktoriia Khoruzhaia , Ekaterina Eroshenko , Ekaterina Derevyanka , Dmitry Yudin

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal development in the clip. Our end-to-end encoder-decoder video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 Zohreh Ghaderi , Leonard Salewski , Hendrik P. A. Lensch

The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility and independence.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yifan Xu , Vineet Kamat , Carol Menassa

In this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply understood as fully…

Machine Learning · Computer Science 2017-02-21 Anurag Kumar , Bhiksha Raj

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Baoqi Pei , Yifei Huang , Jilan Xu , Yuping He , Guo Chen , Fei Wu , Yu Qiao , Jiangmiao Pang

Object SLAM uses additional semantic information to detect and map objects in the scene, in order to improve the system's perception and map representation capabilities. Quadrics and cubes are often used to represent objects, but their…

Robotics · Computer Science 2022-09-23 Xiao Han , Lu Yang