The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies
Abstract
The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model differences significant. DeepSeek and Qwen are the most likely to recommend escalatory action using nuclear weapons; GPT and ERNIE are the least likely. Llama exhibits a distinct bias for action, favoring force, intervention, and cooperation across domains. Inter-rater reliability metrics (Krippendorff's and quadratically weighted Fleiss' ) reveal Llama and ERNIE are the most consistent across runs, with either DeepSeek or GLM the least depending on the domain. We also present a deeper exploration of our scenario variants: (i)~country-level biases tend to exist and vary by model, with country covariates like adversary trade ties and escalation propensity producing weak correlations; (ii)~existential phrasing effects are significant and heterogeneous; (iii)~these country biases interact with phrasing. Overall, the distributions of responses related to the scenarios in our benchmark vary significantly by model, country, and phrasing.
Cite
@article{arxiv.2608.05180,
title = {The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies},
author = {Benjamin Jensen and Ian Reynolds and Yasir Atalan and Martin Pollack and Austin Woo and Robert Sincero},
journal= {arXiv preprint arXiv:2608.05180},
year = {2026}
}
Comments
34 pages, 6 figures