中文

Xmodel-LM技术报告

计算与语言 2024-11-20 v5 人工智能

摘要

我们介绍了Xmodel-LM,一个紧凑高效的1.1B语言模型,在约2万亿个token上进行了预训练。该模型在我们自建的、基于下游任务优化平衡中英文语料的Xdata数据集上训练,尽管规模较小,但表现出卓越的性能。它显著超越了现有类似规模的开源语言模型。我们的模型检查点和代码可在GitHub上公开获取:https://github.com/XiaoduoAILab/XmodelLM。

关键词

引用

@article{arxiv.2406.02856,
  title  = {Xmodel-LM Technical Report},
  author = {Yichuan Wang and Yang Liu and Yu Yan and Qun Wang and Xucheng Huang and Ling Jiang},
  journal= {arXiv preprint arXiv:2406.02856},
  year   = {2024}
}