English

Identifying Reliable Evaluation Metrics for Scientific Text Revision

Computation and Language 2026-01-26 v5

Abstract

Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningful improvements. In this work, we analyse and identify the limitations of these metrics and explore alternative evaluation methods that better align with human judgments. We first conduct a manual annotation study to assess the quality of different revisions. Then, we investigate reference-free evaluation metrics from related NLP domains. Additionally, we examine LLM-as-a-judge approaches, analysing their ability to assess revisions with and without a gold reference. Our results show that LLMs effectively assess instruction-following but struggle with correctness, while domain-specific metrics provide complementary insights. We find that a hybrid approach combining LLM-as-a-judge evaluation and task-specific metrics offers the most reliable assessment of revision quality.

Keywords

Cite

@article{arxiv.2506.04772,
  title  = {Identifying Reliable Evaluation Metrics for Scientific Text Revision},
  author = {Léane Jourdan and Florian Boudin and Richard Dufour and Nicolas Hernandez},
  journal= {arXiv preprint arXiv:2506.04772},
  year   = {2026}
}

Comments

V5 contains the English version, (ACL 2025 main, 26 pages) and V4 contains the French version (TALN 2025, 32 pages), both with corrected results for cramer's v and pairwise accuracy

R2 v1 2026-07-01T03:00:56.247Z