Xmodel-LM技术报告
计算与语言
2024-11-20 v5 人工智能
摘要
我们介绍了Xmodel-LM,一个紧凑高效的1.1B语言模型,在约2万亿个token上进行了预训练。该模型在我们自建的、基于下游任务优化平衡中英文语料的Xdata数据集上训练,尽管规模较小,但表现出卓越的性能。它显著超越了现有类似规模的开源语言模型。我们的模型检查点和代码可在GitHub上公开获取:https://github.com/XiaoduoAILab/XmodelLM。
引用
@article{arxiv.2406.02856,
title = {Xmodel-LM Technical Report},
author = {Yichuan Wang and Yang Liu and Yu Yan and Qun Wang and Xucheng Huang and Ling Jiang},
journal= {arXiv preprint arXiv:2406.02856},
year = {2024}
}