中文

超越通用方案:反向学习用于高度有效的自然语言生成评估提示

计算与语言 2025-09-11 v3

摘要

评估自然语言生成系统具有挑战性 due to valid outputs 的多样性。虽然人类评估是黄金标准,但因而存在不一致性、缺乏标准化和人口统计偏见,限制了可重复性。LLM-based 评估器提供了可扩展的替代方案,但对提示设计高度敏感,其中细微变化可导致显著差异。本文提出一种反向学习方法,从模型输出反向映射到其输入指令, enabling automatic generation of highly effective, model-specific evaluation prompts。该方法仅需一个 evaluation sample,无需耗时的 manual prompt engineering,从而提高了效率和鲁棒性。我们的工作 contribute toward a new direction for more robust and efficient LLM-based evaluation。

关键词

引用

@article{arxiv.2504.21117,
  title  = {Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts},
  author = {Hanhua Hong and Chenghao Xiao and Yang Wang and Yiqi Liu and Wenge Rong and Chenghua Lin},
  journal= {arXiv preprint arXiv:2504.21117},
  year   = {2025}
}

备注

11 pages, accepted by Transactions of the Association for Computational Linguistics (TACL)