当 Flores Bloomz 出错时:机器翻译评估中的跨方向污染
计算与语言
2026-01-29 v1
摘要
大语言模型(LLMs)可能受基准污染影响,导致得分虚高,掩盖记忆作为泛化的现象。在多语言设置下,这种记忆甚至可能转移到“未受污染”语言中。以 FLORES-200 翻译基准作为诊断工具,我们研究了两种 7-8B 参数的指令调优多语言 LLMs:Bloomz,其在 FLORES 上进行训练;以及作为 uncontaminated 对照的 Llama。我们确认 Bloomz 的 FLORES 污染,并演示了机器翻译污染可以是跨方向的,由于目标侧记忆人为地提升了未见翻译方向的性能。进一步分析表明,尽管进行了各种源侧扰动(如改写和实体替换),记忆参考的召回率仍然持续存在。然而,用实体替换会导致 BLEU 持续下降,这表明这是一种有效的记忆在受污染模型中进行探针的方法。
关键词
引用
@article{arxiv.2601.20858,
title = {When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation},
author = {David Tan and Pinzhen Chen and Josef van Genabith and Koel Dutta Chowdhury},
journal= {arXiv preprint arXiv:2601.20858},
year = {2026}
}
备注
5 pages of content, 15 total. 5 figures, 12 tables total. Accepted to EACL 2026 main conference. Code can be found here: github.com/Mr-Ao-25/cross-ling-contamination