English

LLM Scheming Inversely Scales with Pretraining Language Coverage

Artificial Intelligence 2026-06-09 v1 Computation and Language

Abstract

With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2\% higher scores compared to high-resource languages on a five-category scheming index. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.

Cite

@article{arxiv.2607.24769,
  title  = {LLM Scheming Inversely Scales with Pretraining Language Coverage},
  author = {Nathan Truong and Aryan Panda and Rayming Ye and Zoe Sun and Maheep Chaudhary},
  journal= {arXiv preprint arXiv:2607.24769},
  year   = {2026}
}