English

Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation

Computation and Language 2025-09-25 v1 Artificial Intelligence

Abstract

LLM-as-Judge frameworks are increasingly popular for AI evaluation, yet research findings on the relationship between models' generation and judgment abilities remain inconsistent. We investigate this relationship through systematic dataset- and instance-level analyses across 11 models and 21 diverse tasks. Despite both capabilities relying on the same underlying knowledge, our analyses reveal they are only weakly correlated, primarily due to LLMs' sensitivity to the responses being judged. To address this, we propose a self-reference-guided evaluation strategy that leverages a model's own answers as references. This approach significantly strengthens the correlation between generation and judgment abilities, offering a practical path to align these skills and providing a reliable proxy for model selection in evaluation tasks.

Keywords

Cite

@article{arxiv.2509.19880,
  title  = {Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation},
  author = {Wei-Hsiang Lin and Sheng-Lun Wei and Hen-Hsen Huang and Hsin-Hsi Chen},
  journal= {arXiv preprint arXiv:2509.19880},
  year   = {2025}
}

Comments

Accepted as a long findings paper at EMNLP 2025

R2 v1 2026-07-01T05:53:45.206Z