清晰、有说服力的论证:论 frontier AI 安全案例的基础重构
摘要
本文 CONTRIBUTS to frontier AI 系统的 safety cases 初步论争。Safety cases 是针对系统在特定情境下是否安全可接受进行构建、可辩护论证的结构化方法。历史上,这类论证在航空、核能或汽车等安全关键行业中被广泛使用。因此,frontier AI 的 safety cases 正日益受到关注,既在领先 frontier AI 开发者的 safety policy 中,也在来自生成式 AI 领导者提出的国际研究议程中得到体现,如新加坡共识(Singapore Consensus on Global AI Safety Research Priorities)及国际 AI 安全报告。本文对此进行评估。我们指出,alignment 社区内 drawing explicitly on assurance 社区经验的研究存在显著局限。因此,我们旨在重新思考现有的 alignment safety cases 方法。我们整合 safety assurance 领域的现有方法论,阐述其局限性。基于此,我们构建了一个针对 Deceptive Alignment 和 CBRN capabilities 的 safety case 案例研究,参考了 alignment safety case 社区已有的理论 safety case "sketches"。总体而言,我们通过严谨的理论与方法论,引入 safety assurance 领域的整体性见解,以构建更稳健、可辩护且实用的 safety case 方法论框架,帮助保障 frontier AI 系统的安全。
引用
@article{arxiv.2603.08760,
title = {Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases},
author = {Shaun Feakins and Ibrahim Habli and Phillip Morgan},
journal= {arXiv preprint arXiv:2603.08760},
year = {2026}
}
备注
Full paper presented at the International Association of Safe and Ethical AI 2026 Conference (IASEAI 26). 10 pages, 8 figures