English

ToddlerBERTa: Exploiting BabyBERTa for Grammar Learning and Language Understanding

Computation and Language 2023-11-09 v2 Machine Learning

Abstract

We present ToddlerBERTa, a BabyBERTa-like language model, exploring its capabilities through five different models with varied hyperparameters. Evaluating on BLiMP, SuperGLUE, MSGS, and a Supplement benchmark from the BabyLM challenge, we find that smaller models can excel in specific tasks, while larger models perform well with substantial data. Despite training on a smaller dataset, ToddlerBERTa demonstrates commendable performance, rivalling the state-of-the-art RoBERTa-base. The model showcases robust language understanding, even with single-sentence pretraining, and competes with baselines that leverage broader contextual information. Our work provides insights into hyperparameter choices, and data utilization, contributing to the advancement of language models.

Keywords

Cite

@article{arxiv.2308.16336,
  title  = {ToddlerBERTa: Exploiting BabyBERTa for Grammar Learning and Language Understanding},
  author = {Omer Veysel Cagatan},
  journal= {arXiv preprint arXiv:2308.16336},
  year   = {2023}
}

Comments

CoNLL--CMCL 2023 Shared Task: The BabyLM Challenge, Camera-Ready Version

R2 v1 2026-06-28T12:08:50.108Z