English

Harmonizing AI Safety Thresholds

Artificial Intelligence 2026-07-17 v1

Abstract

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

Cite

@article{arxiv.2607.16112,
  title  = {Harmonizing AI Safety Thresholds},
  author = {Wilber Sean Anterola and Matthew Ball and Luis F. Lafuerza and Markov Grey},
  journal= {arXiv preprint arXiv:2607.16112},
  year   = {2026}
}