超越通用方案:反向学习用于高度有效的自然语言生成评估提示
计算与语言
2025-09-11 v3
摘要
评估自然语言生成系统具有挑战性 due to valid outputs 的多样性。虽然人类评估是黄金标准,但因而存在不一致性、缺乏标准化和人口统计偏见,限制了可重复性。LLM-based 评估器提供了可扩展的替代方案,但对提示设计高度敏感,其中细微变化可导致显著差异。本文提出一种反向学习方法,从模型输出反向映射到其输入指令, enabling automatic generation of highly effective, model-specific evaluation prompts。该方法仅需一个 evaluation sample,无需耗时的 manual prompt engineering,从而提高了效率和鲁棒性。我们的工作 contribute toward a new direction for more robust and efficient LLM-based evaluation。
引用
@article{arxiv.2504.21117,
title = {Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts},
author = {Hanhua Hong and Chenghao Xiao and Yang Wang and Yiqi Liu and Wenge Rong and Chenghua Lin},
journal= {arXiv preprint arXiv:2504.21117},
year = {2025}
}
备注
11 pages, accepted by Transactions of the Association for Computational Linguistics (TACL)