中文

对抗策略击败超人类围棋 AI

机器学习 2023-07-14 v4 人工智能 密码学与安全 机器学习

摘要

我们通过训练对抗策略攻击最先进的围棋 AI 系统 KataGo,在以超人类设置运行的 KataGo 上取得了超过 97% 的胜率。我们的对手并非通过下好围棋取胜,而是诱使 KataGo 犯下严重失误。我们的攻击可零样本迁移到其他超人类围棋 AI,并且具有可理解性——人类专家无需算法辅助即可实现该攻击,稳定击败超人类 AI。我们攻击所揭示的核心漏洞即使在经过对抗训练以防御该攻击的 KataGo 智能体中依然存在。我们的结果表明,即便是超人类 AI 系统也可能存在令人惊讶的失效模式。示例对局见 https://goattack.far.ai/。

关键词

引用

@article{arxiv.2211.00241,
  title  = {Adversarial Policies Beat Superhuman Go AIs},
  author = {Tony T. Wang and Adam Gleave and Tom Tseng and Kellin Pelrine and Nora Belrose and Joseph Miller and Michael D. Dennis and Yawen Duan and Viktor Pogrebniak and Sergey Levine and Stuart Russell},
  journal= {arXiv preprint arXiv:2211.00241},
  year   = {2023}
}

备注

Accepted to ICML 2023, see paper for changelog