English

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling

Computer Vision and Pattern Recognition 2026-04-22 v2 Artificial Intelligence Multiagent Systems

Abstract

Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inference typically provides; yet, retraining large models is often computationally expensive and data intensive. We present Visual Reasoning Agent (VRA), a training-free agentic visual reasoning framework that orchestrates off-the-shelf large vision-language models (LVLMs) with a large reasoning model (LRM) through an iterative Think-Critique-Act loop for cross-model verification, self-critique, and recursive refinement. On the remote sensing benchmark VRSBench VQA dataset, VRA consistently outperforms multiple standalone LVLM baselines and achieves up to 40.67\% improvement on challenging question types spanning both perception and reasoning tasks. In addition, integrating three LVLMs with VRA improves the overall accuracy of the standalone LVLMs from 52.8% to 78.8%, demonstrating the effectiveness of agentic reasoning with increased inference-time compute.

Keywords

Cite

@article{arxiv.2509.16343,
  title  = {Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling},
  author = {Chung-En Johnny Yu and Brian Jalaian and Nathaniel D. Bastian},
  journal= {arXiv preprint arXiv:2509.16343},
  year   = {2026}
}

Comments

Accepted to MORS 2026 Artificial Intelligence Workshop Proceedings

R2 v1 2026-07-01T05:46:33.241Z