跨语言之旅:多模态大语言模型跨语言一致性基准测试
计算与语言
2025-08-26 v5 人工智能
计算机视觉与模式识别
机器学习
摘要
多模态大语言模型(MLLM)的快速发展显著提升了其现实应用能力。然而,在跨语言环境下实现一致的性能,尤其是在整合文化知识方面,仍然是一个重大挑战。为更好地评估这一问题,我们引入了两个新基准:KnowRecall 和 VisRecall,用于评估 MLLM 的跨语言一致性。KnowRecall 是一个视觉问答基准,旨在衡量 15 种语言中的事实知识一致性,重点关注全球地标相关的文化和历史问题。VisRecall 通过要求模型在不访问图像的情况下用 9 种语言描述地标外观来评估视觉记忆一致性。实验结果表明,包括专有模型在内的最先进 MLLM 仍然难以实现跨语言一致性。这强调了需要更稳健的方法来生成真正多语言且具备文化意识的模型。
引用
@article{arxiv.2505.15075,
title = {Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs},
author = {Hao Wang and Pinzhi Huang and Jihan Yang and Saining Xie and Daisuke Kawahara},
journal= {arXiv preprint arXiv:2505.15075},
year = {2025}
}
备注
The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action