English

Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments

Machine Learning 2026-02-10 v1 Artificial Intelligence Computation and Language

Abstract

Training large language model (LLM) agents for adversarial games is often driven by episodic objectives such as win rate. In long-horizon settings, however, payoffs are shaped by latent strategic externalities that evolve over time, so myopic optimization and variation-based regret analyses can become vacuous even when the dynamics are predictable. To solve this problem, we introduce Implicit Strategic Optimization (ISO), a prediction-aware framework in which each agent forecasts the current strategic context and uses it to update its policy online. ISO combines a Strategic Reward Model (SRM) that estimates the long-run strategic value of actions with iso-grpo, a context-conditioned optimistic learning rule. We prove sublinear contextual regret and equilibrium convergence guarantees whose dominant terms scale with the number of context mispredictions; when prediction errors are bounded, our bounds recover the static-game rates obtained when strategic externalities are known. Experiments in 6-player No-Limit Texas Hold'em and competitive Pokemon show consistent improvements in long-term return over strong LLM and RL baselines, and graceful degradation under controlled prediction noise.

Keywords

Cite

@article{arxiv.2602.08041,
  title  = {Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments},
  author = {Boyang Xia and Weiyou Tian and Qingnan Ren and Jiaqi Huang and Jie Xiao and Shuo Lu and Kai Wang and Lynn Ai and Eric Yang and Bill Shi},
  journal= {arXiv preprint arXiv:2602.08041},
  year   = {2026}
}
R2 v1 2026-07-01T10:26:52.836Z