English

ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining

Computation and Language 2025-10-24 v4 Artificial Intelligence Machine Learning

Abstract

The emergence of open-source large language models (LLMs) has expanded opportunities for enterprise applications; however, many organizations still lack the infrastructure to deploy and maintain large-scale models. As a result, small LLMs (sLLMs) have become a practical alternative despite inherent performance limitations. While Domain Adaptive Continual Pretraining (DACP) has been explored for domain adaptation, its utility in commercial settings remains under-examined. In this study, we validate the effectiveness of a DACP-based recipe across diverse foundation models and service domains, producing DACP-applied sLLMs (ixi-GEN). Through extensive experiments and real-world evaluations, we demonstrate that ixi-GEN models achieve substantial gains in target-domain performance while preserving general capabilities, offering a cost-efficient and scalable solution for enterprise-level deployment.

Keywords

Cite

@article{arxiv.2507.06795,
  title  = {ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining},
  author = {Seonwu Kim and Yohan Na and Kihun Kim and Hanhee Cho and Geun Lim and Mintae Kim and Seongik Park and Ki Hyun Kim and Youngsub Han and Byoung-Ki Jeon},
  journal= {arXiv preprint arXiv:2507.06795},
  year   = {2025}
}

Comments

Accepted at EMNLP 2025 Industry Track

R2 v1 2026-07-01T03:53:06.043Z