English

Equilibrium Residuals Expose Three Regimes of Matrix-Game Strategic Reasoning in Language Models

Machine Learning 2026-05-12 v1

Abstract

Large language models can score well on named game-theory benchmarks while failing on the same strategic computation once semantic cues are removed. We show this gap with procedurally generated zero-sum matrix games: a model that recognizes familiar games drops to 34%, 18%, and 2% success on anonymous 2×22{\times}2, 3×33{\times}3, and 5×55{\times}5 payoff matrices. The benchmark separates semantic recall, learned approximate Nash computation, and an output-interface bottleneck that limits scale. Training only on 2×22{\times}2 and 3×33{\times}3 games, supervised fine-tuning raises unseen 5×55{\times}5--7×77{\times}7 success from 2% to 61%, while exploitability-reward training averages 37% with high seed variance. We prove that the exploitability residual is 22-Lipschitz in payoff perturbations, unlike discontinuous vertex-returning LP equilibrium selectors, explaining why residual training can transfer under payoff shifts even when formatting instability limits mean performance. A dominated-action padding experiment provides causal evidence: trained models solve 3×33{\times}3 games embedded in much larger matrices, while random-padded controls fail and dense 12×1212{\times}12 games remain near failure. Procedural evaluation is therefore necessary for measuring strategic reasoning, and residual rewards expose a real but format-limited route to approximate equilibrium computation.

Keywords

Cite

@article{arxiv.2605.10410,
  title  = {Equilibrium Residuals Expose Three Regimes of Matrix-Game Strategic Reasoning in Language Models},
  author = {Wenhua Nie and Binhan Luo and Zijie Meng and Jyh-Shing Roger Jang and Ching-Wen Ma},
  journal= {arXiv preprint arXiv:2605.10410},
  year   = {2026}
}