English

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

Computer Vision and Pattern Recognition 2026-05-12 v1 Artificial Intelligence

Abstract

As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores genuinely reflect robust visual understanding? We identify a pervasive vulnerability, the \textbf{Cartesian Shortcut}: visual reasoning benchmarks prevalently build on orthogonal grid-based layouts that can be readily discretized into explicit textual coordinates. Models systematically exploit this property, heavily leveraging text-based deductive reasoning to assist visual problem-solving. To systematically dismantle this shortcut, we introduce \textbf{Polaris-Bench}, which re-formulates 53 visual reasoning tasks in Polar coordinate space with paired Cartesian counterparts as reference, while preserving consistent logical constraints and task semantics -- thus fundamentally breaking the orthogonal prior that models exploit. Comprehensive evaluation across 1414 state-of-the-art MLLMs reveals that frontier models achieving 7070--83%83\% on Cartesian layouts collapse to 3131--39%39\% on Polar equivalents, with degradation persisting even under complete logical equivalence. Moreover, reasoning gains observed on Cartesian layouts are severely diminished on Polar equivalents. These findings expose a critical deficiency in current MLLMs: the lack of topology-invariant visual reasoning.

Keywords

Cite

@article{arxiv.2605.09883,
  title  = {The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space},
  author = {Xia Hu and Zhenrui Yue and Brian Potetz and Howard Zhou and Leonidas Guibas and Chun-Ta Lu and Zhicheng Wang},
  journal= {arXiv preprint arXiv:2605.09883},
  year   = {2026}
}
R2 v1 2026-07-22T07:02:59.976Z