English

Toward Global Large Language Models in Medicine

Computation and Language 2026-01-06 v1

Abstract

Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed the landscape of medicine and holds promise for improving health care quality and expanding access to medical information globally. However, existing LLMs are primarily trained on high-resource languages, limiting their applicability in global medical scenarios. To address this gap, we constructed GlobMed, a large multilingual medical dataset, containing over 500,000 entries spanning 12 languages, including four low-resource languages. Building on this, we established GlobMed-Bench, which systematically assesses 56 state-of-the-art proprietary and open-weight LLMs across multiple multilingual medical tasks, revealing significant performance disparities across languages, particularly for low-resource languages. Additionally, we introduced GlobMed-LLMs, a suite of multilingual medical LLMs trained on GlobMed, with parameters ranging from 1.7B to 8B. GlobMed-LLMs achieved an average performance improvement of over 40% relative to baseline models, with a more than threefold increase in performance on low-resource languages. Together, these resources provide an important foundation for advancing the equitable development and application of LLMs globally, enabling broader language communities to benefit from technological advances.

Keywords

Cite

@article{arxiv.2601.02186,
  title  = {Toward Global Large Language Models in Medicine},
  author = {Rui Yang and Huitao Li and Weihao Xuan and Heli Qi and Xin Li and Kunyu Yu and Yingjian Chen and Rongrong Wang and Jacques Behmoaras and Tianxi Cai and Bibhas Chakraborty and Qingyu Chen and Lionel Tim-Ee Cheng and Marie-Louise Damwanza and Chido Dzinotyiwei and Aosong Feng and Chuan Hong and Yusuke Iwasawa and Yuhe Ke and Linah Kitala and Taehoon Ko and Jisan Lee and Irene Li and Jonathan Chong Kai Liew and Hongfang Liu and Lian Leng Low and Edison Marrese-Taylor and Yutaka Matsuo and Isheanesu Misi and Yilin Ning and Jasmine Chiat Ling Ong and Marcus Eng Hock Ong and Enrico Petretto and Hossein Rouhizadeh and Abiram Sandralegar and Oren Schreier and Iain Bee Huat Tan and Patrick Tan and Daniel Shu Wei Ting and Junjue Wang and Chunhua Weng and Matthew Yu Heng Wong and Fang Wu and Yunze Xiao and Xuhai Xu and Qingcheng Zeng and Zhuo Zheng and Yifan Peng and Douglas Teodoro and Nan Liu},
  journal= {arXiv preprint arXiv:2601.02186},
  year   = {2026}
}

Comments

182 pages, 65 figures

R2 v1 2026-07-01T08:51:01.066Z