English

AIMIP Phase 1: systematic evaluations of AI weather and climate models

Atmospheric and Oceanic Physics 2026-05-20 v2

Abstract

We present the AI weather and climate model intercomparison project (AIMIP), phase 1. Drawing from the rich tradition of intercomparisons in climate model development, we specify a common experiment, output data format, and training constraints (namely, training against historical reanalysis data) for AIMIP Phase 1 models. We aim to identify differences in modeling frameworks and AI architectural choices that influence model behavior, and build trust in AI weather and climate models through open data and evaluation. AIMIP Phase 1 models must simulate the atmosphere given specified historical sea surface temperatures over 1979-2024. We evaluate the models' performance using five major evaluation criteria: biases, trends, response to El Ni\~{n}o-related sea surface temperature anomalies, temporal variability, and out-of-sample generalization tests. We find that the AI models are able to simulate the historical climate and response to forcing as well as a conventional physically-based model, but some AI models underestimate historical warming trends, and their predictions diverge in the out-of-sample generalization tests. We describe the AIMIP Phase 1 dataset that is publicly available for additional evaluations.

Keywords

Cite

@article{arxiv.2605.06944,
  title  = {AIMIP Phase 1: systematic evaluations of AI weather and climate models},
  author = {Brian Henn and Christopher S. Bretherton and Nikolay Koldunov and Christian Lessig and Maria J. Molina and Troy Arcomano and Oliver Watt-Meyer and Guillaume Couairon and Renu Singh and Robert Brunstein and Yana Hasson and Antonia Jost and Noah Brenowitz and Peter Manshausen and Nathaniel Cresswell-Clay and Dale Durran and Kyle Joseph Chen Hall and Janni Yuval and Dmitrii Kochkov and Stephan Hoyer and Ignacio Lopez-Gomez},
  journal= {arXiv preprint arXiv:2605.06944},
  year   = {2026}
}

Comments

48 pages, 25 figures