MCQG-SRefine:基于迭代自我批判、纠正与比较反馈的多选题生成
计算与语言
2025-02-11 v4 人工智能
摘要
自动问题生成(QG)是AI和NLP领域的关键任务,尤其在智能辅导、对话系统和事实验证方面具有重要意义。为专业考试(如美国医学执照考试USMLE)生成多选题尤其具有挑战性,需要专业知识和复杂的多跳推理才能制造高质量问题。然而,当前大语言模型(LLM)如GPT-4在专业MCQG方面因知识过时、幻觉问题以及提示敏感性,导致质量不佳且困难。为此,我们提出MCQG-SRefine,一种基于LLM自我精炼(批判与纠正)框架,将医学病例转换为高质量USMLE风格问题。通过整合专家驱动的提示工程与迭代自我批判和自我纠正反馈,MCQG-SRefine显著提升了人类专家对问题质量和难度的满意度。此外,我们引入一种基于LLM评判的自动化指标,以取代复杂且昂贵的专家评估过程,确保可靠且符合专家标准的评估。
引用
@article{arxiv.2410.13191,
title = {MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback},
author = {Zonghai Yao and Aditya Parashar and Huixue Zhou and Won Seok Jang and Feiyun Ouyang and Zhichao Yang and Hong Yu},
journal= {arXiv preprint arXiv:2410.13191},
year = {2025}
}
备注
Equal contribution for the first two authors. To appear in proceedings of the Main Conference on 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL). Keywords: Question Generation, USMLE, Self-Refine, Self-Critique, and Self-Correction, LLM-as-Judge, AI for Medical Education