English

Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages

Computation and Language 2026-03-25 v1

Abstract

Language barriers affect 27.3 million U.S. residents with non-English language preference, yet professional medical translation remains costly and often unavailable. We evaluated four frontier large language models (GPT-5.1, Claude Opus 4.5, Gemini 3 Pro, Kimi K2) translating 22 medical documents into 8 languages spanning high-resource (Spanish, Chinese, Russian, Vietnamese), medium-resource (Korean, Arabic), and low-resource (Tagalog, Haitian Creole) categories using a five-layer validation framework. Across 704 translation pairs, all models achieved high semantic preservation (LaBSE greater than 0.92), with no significant difference between high- and low-resource languages (p = 0.066). Cross-model back-translation confirmed results were not driven by same-model circularity (delta = -0.0009). Inter-model concordance across four independently trained models was high (LaBSE: 0.946), and lexical borrowing analysis showed no correlation between English term retention and fidelity scores in low-resource languages (rho = +0.018, p = 0.82). These converging results suggest frontier LLMs preserve medical meaning across resource levels, with implications for language access in healthcare.

Keywords

Cite

@article{arxiv.2603.22642,
  title  = {Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages},
  author = {Chukwuebuka Anyaegbuna and Eduardo Juan Perez Guerrero and Jerry Liu and Timothy Keyes and April Liang and Natasha Steele and Stephen Ma and Jonathan Chen and Kevin Schulman},
  journal= {arXiv preprint arXiv:2603.22642},
  year   = {2026}
}

Comments

32 references, 5 tables, 2 figures