English

Spatial-Conditioned Reasoning in Long-Egocentric Videos

Computer Vision and Pattern Recognition 2026-04-09 v2

Abstract

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and short-video reasoning, their spatial reasoning capability in long egocentric sequences remains limited. In this work, we study how explicit spatial signals influence VLM-based video understanding without modifying model architectures or inference procedures. We introduce Sanpo-D, a fine-grained re-annotation of the Google Sanpo dataset, and benchmark multiple VLMs on navigation-oriented spatial queries. To examine input-level inductive bias, we further fuse depth maps with RGB frames and evaluate their impact on spatial reasoning. Our results reveal a trade-off between general-purpose accuracy and spatial specialization, showing that depth-aware and spatially grounded representations can improve performance on safety-critical tasks such as pedestrian and obstruction detection.

Keywords

Cite

@article{arxiv.2601.18100,
  title  = {Spatial-Conditioned Reasoning in Long-Egocentric Videos},
  author = {James Tribble and Hao Wang and Si-En Hong and Chaoyi Zhou and Ashish Bastola and Siyu Huang and Abolfazl Razi},
  journal= {arXiv preprint arXiv:2601.18100},
  year   = {2026}
}
R2 v1 2026-07-01T09:19:36.559Z