English

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

Computation and Language 2026-07-12 v1

Abstract

This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found in https://chinese-babylm.github.io/.

Cite

@article{arxiv.2607.10745,
  title  = {The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese},
  author = {Siyuan Song and Zhiheng Qian and Yunhao Zhang and Linyang He and Xiaozhe Ji and Yingxin Lin and Hongao Zhu and Chongtian Shao and Chuhan Lang and Luan Li and Rui Wang and Renfen Hu and Shaonan Wang and Hai Hu},
  journal= {arXiv preprint arXiv:2607.10745},
  year   = {2026}
}

Comments

8 pages, 4 tables; work in progress

R2 v1 2026-07-22T20:36:26.507Z