English

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design

Computer Vision and Pattern Recognition 2026-05-12 v3 Artificial Intelligence Machine Learning

Abstract

Visual perception in modern Vision-Language Models (VLMs) is constrained by a perceptual bandwidth bottleneck: a broad field of view preserves global context but sacrifices the fine-grained details required for complex reasoning. We argue that high-resolution visual reasoning is therefore not only semantic reasoning but also task-relevant evidence acquisition under limited perceptual bandwidth. Inspired by active vision and information foraging, we formalise this process as sequential Bayesian optimal experimental design (S-BOED), where an agent decides which visual evidence to acquire before answering. Since exact Bayesian inference is intractable in continuous gigapixel spaces, we derive a tractable coverage--resolution objective as a proxy for task-relevant information gain. We instantiate this framework with FOVEA, a training-free procedure that refines VLM crop proposals through evidence-oriented probing. Experiments on high-resolution benchmarks show consistent gains over direct and ReAct-style baselines, with particularly strong improvements in search-dominated remote-sensing settings.

Keywords

Cite

@article{arxiv.2605.01345,
  title  = {The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design},
  author = {Anjie Liu and Ziqin Gong and Yan Song and Yuxiang Chen and Xiaolong Liu and Hengtong Lu and Kaike Zhang and Chen Wei and Jun Wang},
  journal= {arXiv preprint arXiv:2605.01345},
  year   = {2026}
}

Comments

27 pages, 5 figures, accepted at ICML 2026

R2 v1 2026-07-01T12:46:30.263Z