English

Pretraining on the Test Set Is All You Need

Computation and Language 2023-09-19 v1 Artificial Intelligence

Abstract

Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM \textbf{phi-CTNL} (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. \textbf{phi-CTNL} also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.

Keywords

Cite

@article{arxiv.2309.08632,
  title  = {Pretraining on the Test Set Is All You Need},
  author = {Rylan Schaeffer},
  journal= {arXiv preprint arXiv:2309.08632},
  year   = {2023}
}

Comments

3 pages, satire

R2 v1 2026-06-28T12:22:57.560Z