关于评估临床问题中 LLM 输出的课程共享任务
计算与语言
2024-08-02 v1
摘要
本文介绍了我们在 2023/2024 年于德国达尔茨堡技术大学举办的 Foundations of Language Technology (FoLT) 课程期间组织的共享任务,旨在评估大型语言模型 (LLM) 在生成针对与健康相关临床问题的有害答案方面的表现。我们描述了任务设计考量,并报告了来自学生的反馈。我们期望该任务以及本文所报告的发现对教师教授自然语言处理 (NLP) 及设计课程作业的人群具有相关性。
引用
@article{arxiv.2408.00122,
title = {A Course Shared Task on Evaluating LLM Output for Clinical Questions},
author = {Yufang Hou and Thy Thy Tran and Doan Nam Long Vu and Yiwen Cao and Kai Li and Lukas Rohde and Iryna Gurevych},
journal= {arXiv preprint arXiv:2408.00122},
year = {2024}
}
备注
accepted at the sixth Workshop on Teaching NLP (co-located with ACL 2024)