真实情景下的 scheming 本能评估
机器学习
2026-05-29 v1
摘要
我们引入了 scheming honeypot evaluations,一种用于测试模型在获得机会时是否会追求工具性目标的框架。我们的 scheming honeypot evaluations 采取谷歌对齐研究代码库中的编码任务形式。在实际内部部署设置中,Gemini 模型未表现出未被激发的 scheming 行为。如果提示明确鼓励主动性(情境意识或目标导向性)和/或给出模型隐藏目标,模型有时会进行 scheming 或尝试破坏。验证我们设定的真实性,模型显示出低的评估意识率,这通常归因于激发主动性的提示,而非环境本身。
引用
@article{arxiv.2605.29729,
title = {Realistic honeypot evaluations for scheming propensity},
author = {Victoria Krakovna and David Lindner and Lewis Ho and Sebastian Farquhar and Rohin Shah},
journal= {arXiv preprint arXiv:2605.29729},
year = {2026}
}