中文

评估大语言模型在泰语方言中的表现:自动基准测试与人类评估

计算与语言 2025-04-09 v1

摘要

大型语言模型在 various NLP 任务中显示出鼓舞人心的结果。尽管如此,针对 underrepresented 语言(尤其是当地方言)的鲁棒性和一致性仍基本未被探索,现有基准测试也关注主要方言,忽略了大型语言模型在 local dialect 文本方面的能力。本文引入了一个覆盖泰语北部 (Lanna)、东北部 (Isan) 和南部 (Dambro) 方言的泰语 local dialect 基准测试,评估大型语言模型在五个 NLP 任务上的表现:摘要、问答、翻译、对话和食物相关任务。此外,我们提出了一个用于评估泰语 local dialect 生成流畅度和方言特定准确性的人类评估指南和指标。结果显示,在 local Thai 方言中,大型语言模型的表现显著下降,仅有专有模型如 GPT-4o 和 Gemini2 能表现出一些流畅性。

关键词

引用

@article{arxiv.2504.05898,
  title  = {Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation},
  author = {Peerat Limkonchotiwat and Kanruethai Masuk and Surapon Nonesung and Chalermpun Mai-On and Sarana Nutanong and Wuttikorn Ponwitayarat and Potsawee Manakul},
  journal= {arXiv preprint arXiv:2504.05898},
  year   = {2025}
}

备注

Datasets and codes are available at https://github.com/mrpeerat/Thai_local_benchmark