English

VinaLLaMA: LLaMA-based Vietnamese Foundation Model

Computation and Language 2023-12-19 v1

Abstract

In this technical report, we present VinaLLaMA, an open-weight, state-of-the-art (SOTA) Large Language Model for the Vietnamese language, built upon LLaMA-2 with an additional 800 billion trained tokens. VinaLLaMA not only demonstrates fluency in Vietnamese but also exhibits a profound understanding of Vietnamese culture, making it a truly indigenous model. VinaLLaMA-7B-chat, trained on 1 million high-quality synthetic samples, achieves SOTA results on key benchmarks, including VLSP, VMLU, and Vicuna Benchmark Vietnamese, marking a significant advancement in the Vietnamese AI landscape and offering a versatile resource for various applications.

Cite

@article{arxiv.2312.11011,
  title  = {VinaLLaMA: LLaMA-based Vietnamese Foundation Model},
  author = {Quan Nguyen and Huy Pham and Dung Dao},
  journal= {arXiv preprint arXiv:2312.11011},
  year   = {2023}
}

Comments

VinaLLaMA Technical Report - 13 pages

R2 v1 2026-06-28T13:54:21.769Z