English

Enriching Biomedical Knowledge for Low-resource Language Through Large-Scale Translation

Computation and Language 2023-01-31 v3 Artificial Intelligence

Abstract

Biomedical data and benchmarks are highly valuable yet very limited in low-resource languages other than English such as Vietnamese. In this paper, we make use of a state-of-the-art translation model in English-Vietnamese to translate and produce both pretrained as well as supervised data in the biomedical domains. Thanks to such large-scale translation, we introduce ViPubmedT5, a pretrained Encoder-Decoder Transformer model trained on 20 million translated abstracts from the high-quality public PubMed corpus. ViPubMedT5 demonstrates state-of-the-art results on two different biomedical benchmarks in summarization and acronym disambiguation. Further, we release ViMedNLI - a new NLP task in Vietnamese translated from MedNLI using the recently public En-vi translation model and carefully refined by human experts, with evaluations of existing methods against ViPubmedT5.

Keywords

Cite

@article{arxiv.2210.05598,
  title  = {Enriching Biomedical Knowledge for Low-resource Language Through Large-Scale Translation},
  author = {Long Phan and Tai Dang and Hieu Tran and Trieu H. Trinh and Vy Phan and Lam D. Chau and Minh-Thang Luong},
  journal= {arXiv preprint arXiv:2210.05598},
  year   = {2023}
}
R2 v1 2026-06-28T03:16:04.508Z