English

From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring

Computation and Language 2026-03-09 v1

Abstract

Large language models (LLMs) have recently reshaped Automated Essay Scoring (AES), yet prior studies typically examine individual techniques in isolation, limiting understanding of their relative merits for English as a Second Language (L2) writing. To bridge this gap, we presents a comprehensive comparison of major LLM-based AES paradigms on IELTS Writing Task~2. On this unified benchmark, we evaluate four approaches: (i) encoder-based classification fine-tuning, (ii) zero- and few-shot prompting, (iii) instruction tuning and Retrieval-Augmented Generation (RAG), and (iv) Supervised Fine-Tuning combined with Direct Preference Optimization (DPO) and RAG. Our results reveal clear accuracy-cost-robustness trade-offs across methods, the best configuration, integrating k-SFT and RAG, achieves the strongest overall results with F1-Score 93%. This study offers the first unified empirical comparison of modern LLM-based AES strategies for English L2, promising potential in auto-grading writing tasks. Code is public at https://github.com/MinhNguyenDS/LLM_AES-EnL2

Keywords

Cite

@article{arxiv.2603.06424,
  title  = {From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring},
  author = {Minh Hoang Nguyen and Vu Hoang Pham and Xuan Thanh Huynh and Phuc Hong Mai and Vinh The Nguyen and Quang Nhut Huynh and Huy Tien Nguyen and Tung Le},
  journal= {arXiv preprint arXiv:2603.06424},
  year   = {2026}
}

Comments

19 pages, 10 figures, 7 tables

R2 v1 2026-07-01T11:07:12.117Z