English

Architectural Trade-offs in Small Language Models Under Compute Constraints

Computation and Language 2025-12-25 v1 Machine Learning

Abstract

We present a systematic empirical study of small language models under strict compute constraints, analyzing how architectural choices and training budget interact to determine performance. Starting from a linear next-token predictor, we progressively introduce nonlinearities, self-attention, and multi-layer transformer architectures, evaluating each on character-level modeling of Tiny Shakespeare and word-level modeling of Penn Treebank (PTB) and WikiText-2. We compare models using test negative log-likelihood (NLL), parameter count, and approximate training FLOPs to characterize accuracy-efficiency trade-offs. Our results show that attention-based models dominate MLPs in per-FLOP efficiency even at small scale, while increasing depth or context without sufficient optimization can degrade performance. We further examine rotary positional embeddings (RoPE), finding that architectural techniques successful in large language models do not necessarily transfer to small-model regimes.

Keywords

Cite

@article{arxiv.2512.20877,
  title  = {Architectural Trade-offs in Small Language Models Under Compute Constraints},
  author = {Shivraj Singh Bhatti},
  journal= {arXiv preprint arXiv:2512.20877},
  year   = {2025}
}

Comments

15 pages, 11 images

R2 v1 2026-07-01T08:39:27.638Z