面向LLMs医学伦理评估的综合问答基准——MedEthicsQA
计算与语言
2025-07-01 v1 人工智能
摘要
尽管医学大语言模型(MedLLMs)在临床任务中展现出显著潜力,但其伦理安全性仍不足以被探讨。本文介绍 ,一个包含 项多选题和 项开放题的综合基准,用于评估医学伦理。我们系统构建了层次化医学伦理标准体系。该基准涵盖广泛使用的医学数据集、权威题库及来自 PubMed 文献的场景。通过多阶段筛选和多方面专家验证,确保数据集的可靠性,错误率低于 。评估最新医学伦理模型显示,其在回答医学伦理问题方面的表现低于基础模型,揭示了医学伦理对齐的不足。该数据集已注册于 CC BY-NC 4.0 许可证,链接为 https://github.com/JianhuiWei7/MedEthicsQA。
引用
@article{arxiv.2506.22808,
title = {MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs},
author = {Jianhui Wei and Zijie Meng and Zikai Xiao and Tianxiang Hu and Yang Feng and Zhijie Zhou and Jian Wu and Zuozhu Liu},
journal= {arXiv preprint arXiv:2506.22808},
year = {2025}
}
备注
20 pages