中文

探索 使用 LLM 对 编程 教育 中 学生 自我 解释 自动 评估 的 有效性

人机交互 2026-05-22 v1 机器学习

摘要

worked examples 是特定领域中解决问题的逐步解答,旨在帮助学生获取领域特定的解决问题技能。将 worked examples 与 self-explanations 结合——后者要求学生解释而非被动学习每个解决步骤——可增强其效果。主要挑战在于评估学生解释的正确性。在当前方法中,学生的解释被 judged by their semantic similarity to an instructor's or domain expert's explanation. 鉴于最近在 LLM 基础自动评分方面的进展,尚不清楚语义相似性方法是否仍是自动对类似 essay 或代码 解释 的文本学生响应的最有效技术。比较这些方法也需要提供独特特征的数据集,如平衡的 类分布 和 用于 自动 评分 任务的 领域特定 标记 数据。在本文中,我们提出了一种严格的比较, 将 LLM 与 语义相似性用于 自动 评分, 其 被 视为 二元 分类 任务。

关键词

引用

@article{arxiv.2605.21614,
  title  = {Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education},
  author = {Arun-Balajiee Lekshmi-Narayanan and Mohammad Hassany and Peter Brusilovsky},
  journal= {arXiv preprint arXiv:2605.21614},
  year   = {2026}
}