English

Evaluating LLMs on Chinese Idiom Translation

Computation and Language 2025-08-15 v1

Abstract

Idioms, whose figurative meanings usually differ from their literal interpretations, are common in everyday language, especially in Chinese, where they often contain historical references and follow specific structural patterns. Despite recent progress in machine translation with large language models, little is known about Chinese idiom translation. In this work, we introduce IdiomEval, a framework with a comprehensive error taxonomy for Chinese idiom translation. We annotate 900 translation pairs from nine modern systems, including GPT-4o and Google Translate, across four domains: web, news, Wikipedia, and social media. We find these systems fail at idiom translation, producing incorrect, literal, partial, or even missing translations. The best-performing system, GPT-4, makes errors in 28% of cases. We also find that existing evaluation metrics measure idiom quality poorly with Pearson correlation below 0.48 with human ratings. We thus develop improved models that achieve F1_1 scores of 0.68 for detecting idiom translation errors.

Keywords

Cite

@article{arxiv.2508.10421,
  title  = {Evaluating LLMs on Chinese Idiom Translation},
  author = {Cai Yang and Yao Dou and David Heineman and Xiaofeng Wu and Wei Xu},
  journal= {arXiv preprint arXiv:2508.10421},
  year   = {2025}
}

Comments

Accepted at COLM 2025

R2 v1 2026-07-01T04:49:27.938Z