English

iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision

Human-Computer Interaction 2026-02-16 v1

Abstract

Large Language Models (LLMs) have become indispensable for evaluating writing. However, text feedback they provide is often unintelligible, generic, and not specific to user criteria. Inspired by structured rubrics in education and intelligible AI explanations, we propose iRULER following identified design guidelines to \textit{scaffold} the review process by \textit{specific} criteria, providing \textit{justification} for score selection, and offering \textit{actionable} revisions to target different quality levels. To \textit{qualify} user-defined criteria, we recursively used iRULER with a rubric-of-rubrics to iteratively \textit{refine} rubrics. In controlled experiments on writing revision and rubric creation, iRULER most improved validated LLM-judged review scores and was perceived as most helpful and aligned compared to read-only rubric and text-based LLM feedback. Qualitative findings further support how iRULER satisfies the design guidelines for user-defined feedback. This work contributes interactive rubric tools for intelligible LLM-based review and revision of writing, and user-defined rubric creation.

Keywords

Cite

@article{arxiv.2602.12779,
  title  = {iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision},
  author = {Jingwen Bai and Wei Soon Cheong and Philippe Muller and Brian Y Lim},
  journal= {arXiv preprint arXiv:2602.12779},
  year   = {2026}
}

Comments

To Appear at CHI 2026

R2 v1 2026-07-01T10:35:05.653Z