中文

AI 法规评估基准:面向 NLP 与 RAG 系统的开放、透明且可复现的评估数据集

人工智能 2026-03-11 v1

摘要

AI 在 heterogeneous public and societal sectors 的 rapid 推广导致对 regulatory standards 与 frameworks 的 compliance 需求日益增加。欧盟 AI 法规(EU AI Act)已成为监管格局中的里程碑。开发能够 elicit AI systems 符合此类 standards 的 solutions 常常受限于 lack of resources,限制了其 semi-automated 或 automated evaluation 能力。这导致依赖 manual work,容易 error-prone,resource-limited,或 limited to cases not clearly described by regulation。本 paper 提出一种 open, transparent, reproducible 的 method,用于 creating 促进 NLP models 评估的 resource,尤其 focus on RAG systems。我们 developed 一个 dataset,包含 risk-level classification, article retrieval, obligation generation, and question-answering 四类 tasks 用于 EU AI Act。dataset 文件采用 machine-to-machine appropriate format。为生成文件,我们 utilise domain knowledge 作为 exegetical basis,结合 large language models 的 processing 与 reasoning power 生成 scenarios 及 respective tasks。Our methodology 演示了一种利用 language models 实现 grounded generation 且 document relevancy high 的方法。此外,我们克服了诸多局限性,例如 risk-levels decision boundaries not explicitly defined within EU AI Act(如 limited and minimal cases)。最后,我们通过 evaluate 一个 RAG-based solution demonstration dataset effectiveness,其在 prohibited 与 high-risk scenarios 下分别达到 0.87 与 0.85 F1-score。

关键词

引用

@article{arxiv.2603.09435,
  title  = {AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems},
  author = {Athanasios Davvetas and Michael Papademas and Xenia Ziouvelou and Vangelis Karkaletsis},
  journal= {arXiv preprint arXiv:2603.09435},
  year   = {2026}
}

备注

10 pages, 1 figure, 4 tables, 2 equations