中文

冗长LLM输出的影响:翻译评估案例研究

计算与语言 2024-10-02 v1

摘要

本文探讨了冗长LLM翻译输出对评估的影响。我们首先展示了这种行为在来自WMT 2024通用共享任务的多个LLM输出中出现的普遍性。随后,我们确定引发冗长输出的主要触发因素,包括安全、版权 concerns以及短输入查询缺乏上下文。最后,我们表明,忽视这一行为会对更冗长的LLM在自动评估和人类评估中的表现不公平,凸显了为更准确的未来评估而需要解决这一问题的必要性。

关键词

引用

@article{arxiv.2410.00863,
  title  = {On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation},
  author = {Eleftheria Briakou and Zhongtao Liu and Colin Cherry and Markus Freitag},
  journal= {arXiv preprint arXiv:2410.00863},
  year   = {2024}
}