中文

探索生成式语言模型用于长作文自动评分中的摘要化方法

计算与语言 2025-11-20 v4 机器学习

摘要

BERT及其变体广泛应用于自动评分。然而,这些基于编码器的模型在512个标记的限制上存在不足,导致长作文自动评分效果受限。因此,本研究探索通过摘要化和提示技术,使用生成式语言模型进行长作文自动评分。结果表明,评分准确率显著提升,在Learning Agency Lab Automated Essay Scoring 2.0数据集上,QWK从0.822提升至0.8878。

关键词

引用

@article{arxiv.2510.22830,
  title  = {Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays},
  author = {Haowei Hua and Hong Jiao and Xinyi Wang},
  journal= {arXiv preprint arXiv:2510.22830},
  year   = {2025}
}

备注

19 pages, 5 Tables 7 Figures, Presentation at Artificial Intelligence in Measurement and Education Conference (AIME-Con)