Back to Research papers
Research paper index

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

Siyuan Song, Zhiheng Qian, Yunhao Zhang, Linyang He, Xiaozhe Ji, Yingxin Lin, Hongao Zhu, Chongtian Shao, Chuhan Lang, Luan Li, Rui Wang, Renfen Hu, Shaonan Wang, Hai Hu

arXiv:2607.10745Published July 12, 20260 citations
  • cs.CL

Abstract

This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found in https://chinese-babylm.github.io/.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.