探索生成式语言模型用于长作文自动评分中的摘要化方法
计算与语言
2025-11-20 v4 机器学习
摘要
BERT及其变体广泛应用于自动评分。然而,这些基于编码器的模型在512个标记的限制上存在不足,导致长作文自动评分效果受限。因此,本研究探索通过摘要化和提示技术,使用生成式语言模型进行长作文自动评分。结果表明,评分准确率显著提升,在Learning Agency Lab Automated Essay Scoring 2.0数据集上,QWK从0.822提升至0.8878。
引用
@article{arxiv.2510.22830,
title = {Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays},
author = {Haowei Hua and Hong Jiao and Xinyi Wang},
journal= {arXiv preprint arXiv:2510.22830},
year = {2025}
}
备注
19 pages, 5 Tables 7 Figures, Presentation at Artificial Intelligence in Measurement and Education Conference (AIME-Con)