English

Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs

Machine Learning 2025-08-26 v1 Artificial Intelligence Information Retrieval

Abstract

How does retrieval performance scale with pretraining FLOPs? We benchmark retrieval performance across LLM model sizes from 125 million parameters to 7 billion parameters pretrained on datasets ranging from 1 billion tokens to more than 2 trillion tokens. We find that retrieval performance on zero-shot BEIR tasks predictably scales with LLM size, training duration, and estimated FLOPs. We also show that In-Context Learning scores are strongly correlated with retrieval scores across retrieval tasks. Finally, we highlight the implications this has for the development of LLM-based retrievers.

Keywords

Cite

@article{arxiv.2508.17400,
  title  = {Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs},
  author = {Jacob Portes and Connor Jennings and Erica Ji Yuen and Sasha Doubov and Michael Carbin},
  journal= {arXiv preprint arXiv:2508.17400},
  year   = {2025}
}

Comments

15 pages, 4 figures

R2 v1 2026-07-01T05:03:33.141Z