中文

RAIL 在野外:使用 Anthropic 的价值数据实施负责任 AI 评估

地球与行星天体物理 2025-05-02 v1 太阳与恒星天体物理

摘要

随着 AI 系统嵌入实际应用,确保其符合伦理标准至关重要。虽然现有 AI 伦理框架强调公平性、透明性和可解释性,但往往缺乏可操作的评估方法。本文引入一种系统方法,利用负责任 AI 实验室(RAIL) 框架,其中包括八个可衡量的维度,用于评估大型语言模型(LLM)的规范行为。我们将其应用于 Anthropic 的“Values in the Wild”数据集,该数据集包含超过 308,000 条来自 Claude 的匿名对话以及超过 3,000 条标注的价值表达。我们的研究将这些价值映射到 RAIL 维度,计算合成分数,并提供关于 LLM 在实际应用中伦理行为的见解。

关键词

引用

@article{arxiv.2505.00187,
  title  = {Reduced solar quadrupole moment compensates for lack of asteroids in long-term solar system integrations},
  author = {Richard E. Zeebe and Ilja J. Kocken},
  journal= {arXiv preprint arXiv:2505.00187},
  year   = {2025}
}

备注

Final Accepted Version, The Astronomical Journal