English

Monte Carlo Expected Threat (MOCET) Scoring

Machine Learning 2025-11-24 v1 Artificial Intelligence Human-Computer Interaction

Abstract

Evaluating and measuring AI Safety Level (ASL) threats are crucial for guiding stakeholders to implement safeguards that keep risks within acceptable limits. ASL-3+ models present a unique risk in their ability to uplift novice non-state actors, especially in the realm of biosecurity. Existing evaluation metrics, such as LAB-Bench, BioLP-bench, and WMDP, can reliably assess model uplift and domain knowledge. However, metrics that better contextualize "real-world risks" are needed to inform the safety case for LLMs, along with scalable, open-ended metrics to keep pace with their rapid advancements. To address both gaps, we introduce MOCET, an interpretable and doubly-scalable metric (automatable and open-ended) that can quantify real-world risks.

Keywords

Cite

@article{arxiv.2511.16823,
  title  = {Monte Carlo Expected Threat (MOCET) Scoring},
  author = {Joseph Kim and Saahith Potluri},
  journal= {arXiv preprint arXiv:2511.16823},
  year   = {2025}
}

Comments

Accepted to NeurIPS 2025 BioSafe GenAI

R2 v1 2026-07-01T07:48:06.663Z