English

PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency

Computation and Language 2024-10-23 v2 Artificial Intelligence Machine Learning

Abstract

We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK Normalization and Z-Loss to ensure training stability during the training process. Post-training techniques, including Supervised Fine-Tuning and Direct Preference Optimization, were applied to refine the model's performance. Benchmark evaluations suggest that PLaMo-100B performs well, particularly in Japanese-specific tasks, achieving results that are competitive with frontier models like GPT-4. The base model is available at https://huggingface.co/pfnet/plamo-100b.

Keywords

Cite

@article{arxiv.2410.07563,
  title  = {PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency},
  author = {Preferred Elements and : and Kenshin Abe and Kaizaburo Chubachi and Yasuhiro Fujita and Yuta Hirokawa and Kentaro Imajo and Toshiki Kataoka and Hiroyoshi Komatsu and Hiroaki Mikami and Tsuguo Mogami and Shogo Murai and Kosuke Nakago and Daisuke Nishino and Toru Ogawa and Daisuke Okanohara and Yoshihiko Ozaki and Shotaro Sano and Shuji Suzuki and Tianqi Xu and Toshihiko Yanase},
  journal= {arXiv preprint arXiv:2410.07563},
  year   = {2024}
}
R2 v1 2026-06-28T19:15:33.162Z