MedQA-CS:基于 OSCE 风格的用于评估大语言模型临床技能的基准
人工智能
2026-01-21 v2 计算与语言
摘要
人工智能(AI)和大语言模型(LLMs)在医疗领域需要先进的临床技能(CS),但现有基准未能全面评估这些技能。我们引入 MedQA-CS,一种受医学教育中客观结构临床考试(OSCE)启发的 AI-SCE 框架,以弥补这一差距。MedQA-CS 通过两种指令遵循任务来评估 LLMs:LLM-as-medical-student 和 LLM-as-CS-examiner,旨�反映真实临床情景。我们的贡献包括开发 MedQA-CS,一个包含公开数据和专家标注的全面评估框架,以及对 LLMs 作为临床技能评估者在定性和定量方面的可靠性进行评估。我们的实验表明,MedQA-CS 是比传统多选题 QA 基准(如 MedQA)更具挑战性的临床技能评估基准。结合现有基准,MedQA-CS 实现了对开放式和封闭式 LLMs 临床能力的更全面评估。
引用
@article{arxiv.2410.01553,
title = {MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills},
author = {Zonghai Yao and Zihao Zhang and Chaolong Tang and Xingyu Bian and Youxia Zhao and Zhichao Yang and Junda Wang and Huixue Zhou and Won Seok Jang and Feiyun Ouyang and Hong Yu},
journal= {arXiv preprint arXiv:2410.01553},
year = {2026}
}
备注
To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026