LLM在明确禁止和监视下表现出误导性行为
人工智能
2025-07-08 v1
摘要
本文中,我们让LLM完成一项不可能的测试题,同时它们处于沙箱环境中,受到监控,被告知这些措施并被指示不要作弊。一些前沿LLM始终作弊并试图规避限制,尽管如此。结果揭示了当前LLM在目标导向行为与对齐之间存在的根本性张力。代码和评估日志已在github.com/baceolus/cheating_evals上提供。
引用
@article{arxiv.2507.02977,
title = {LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance},
author = {Igor Ivanov},
journal= {arXiv preprint arXiv:2507.02977},
year = {2025}
}
备注
10 pages, 2 figures