面向 LLM 对抗鲁棒性 assurance 的本体论驱动论证方法
人工智能
2024-10-11 v1
摘要
尽管大型语言模型(LLM)在适应性方面表现突出,但在确保其安全性、透明性和可解释性方面仍面临挑战。鉴于 LLM 对对抗攻击的脆弱性,需不断结合对抗训练和防护措施进行防御。然而,管理持续确保鲁棒性所隐含且异构的知识困难重重。我们提出一种基于形式化论证的 LLM 对抗鲁棒性 assurance 方法。利用本体论进行形式化,我们结构化了最先进的攻击与防御方法,促进了创建人类可读的 assurance 案例以及机器可读的表示。我们以英文语言和代码翻译任务为例演示了其应用,并就理论与实践提供了启示,面向工程师、数据科学家、用户和审计员等不同受众。
引用
@article{arxiv.2410.07962,
title = {Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation},
author = {Tomas Bueno Momcilovic and Beat Buesser and Giulio Zizzo and Mark Purcell and Dian Balta},
journal= {arXiv preprint arXiv:2410.07962},
year = {2024}
}
备注
To be published in xAI 2024, late-breaking track