English

AIMO Interpretability Challenge

Artificial Intelligence 2026-07-15 v1

Abstract

We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms. The challenge is motivated by a central limitation of standard reasoning benchmarks: strong final-answer accuracy does not reveal whether a model relies on stable reasoning mechanisms or exploits brittle reasoning shortcuts. Building on AI Mathematical Olympiad (AIMO) problems and submissions, together with resources from the Fields Model Initiative, the competition will provide (1) newly-published olympiad-level math reasoning problems and their symbolic representations, allowing generation of novel functional variants, (2) access to frontier reasoning models, and (3) assessments of models' adversarial robustness on these problems. Participants will use these resources, along with our computing infrastructure support, to develop methods for identifying which models solve problems robustly. Our competition will also create a new, open robustness benchmark and baseline systems, aiming to provide a lasting foundation for standard benchmarking in mathematical reasoning and interpretability. Scientifically, the competition connects interpretability and generalization research around a central question in AI research: can we determine if, and to what extent, the decision-making of frontier AI models is generalizable and thus, reliable?

Cite

@article{arxiv.2607.13899,
  title  = {AIMO Interpretability Challenge},
  author = {Michal Štefánik and Philipp Mondorf and Andreas Waldis and Qianying Liu and Chuan Yang and Michal Spiegel and Josef Kuchař and Marek Kadlčík and Adam Vawda-Oomerjee and Chaoran Liu and Simon Frieder and Barbara Plank and Fazl Barez and Pontus Stenetorp},
  journal= {arXiv preprint arXiv:2607.13899},
  year   = {2026}
}

Comments

Accepted Competition at NeurIPS 2026