中文

面向LLMs医学伦理评估的综合问答基准——MedEthicsQA

计算与语言 2025-07-01 v1 人工智能

摘要

尽管医学大语言模型(MedLLMs)在临床任务中展现出显著潜力,但其伦理安全性仍不足以被探讨。本文介绍 MedEthicsQA\textbf{MedEthicsQA},一个包含 5,623\textbf{5,623} 项多选题和 5,351\textbf{5,351} 项开放题的综合基准,用于评估医学伦理。我们系统构建了层次化医学伦理标准体系。该基准涵盖广泛使用的医学数据集、权威题库及来自 PubMed 文献的场景。通过多阶段筛选和多方面专家验证,确保数据集的可靠性,错误率低于 2.72%2.72\%。评估最新医学伦理模型显示,其在回答医学伦理问题方面的表现低于基础模型,揭示了医学伦理对齐的不足。该数据集已注册于 CC BY-NC 4.0 许可证,链接为 https://github.com/JianhuiWei7/MedEthicsQA。

关键词

引用

@article{arxiv.2506.22808,
  title  = {MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs},
  author = {Jianhui Wei and Zijie Meng and Zikai Xiao and Tianxiang Hu and Yang Feng and Zhijie Zhou and Jian Wu and Zuozhu Liu},
  journal= {arXiv preprint arXiv:2506.22808},
  year   = {2025}
}

备注

20 pages