English

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

Computer Vision and Pattern Recognition 2026-04-07 v2

Abstract

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently preserves. In this work, we introduce the Panorama-Language Modeling (PLM)paradigm, a unified 360360^\circ vision-language reasoning that is more than the sum of its pinhole counterparts. Besides, we present PanoVQA, a large-scale panoramic VQA dataset that involves adverse omni-scenes, enabling comprehensive reasoning under object occlusions and driving accidents. To establish a foundation for PLM, we develop a plug-and-play panoramic sparse attention module that allows existing pinhole-based VLMs to process equirectangular panoramas without retraining. Extensive experiments demonstrate that our PLM achieves superior robustness and holistic reasoning under challenging omni-scenes, yielding understanding greater than the sum of its narrow parts. Project page: https://github.com/InSAI-Lab/PanoVQA.

Keywords

Cite

@article{arxiv.2603.09573,
  title  = {More than the Sum: Panorama-Language Models for Adverse Omni-Scenes},
  author = {Weijia Fan and Ruiping Liu and Jiale Wei and Yufan Chen and Junwei Zheng and Zichao Zeng and Jiaming Zhang and Qiufu Li and Linlin Shen and Rainer Stiefelhagen},
  journal= {arXiv preprint arXiv:2603.09573},
  year   = {2026}
}

Comments

Accepted by CVPR 2026. Project page: https://github.com/InSAI-Lab/PanoVQA