English

Learn from your own latents and not from tokens: A sample-complexity theory

Machine Learning 2026-05-28 v1

Abstract

Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which networks are trained to predict their \emph{own} latent representations of related views or masked regions, as in data2vec and JEPA -- an idea related to predictive-coding accounts of the cortex. Despite strong empirical results, the theoretical understanding of these methods remains limited. Central questions include: by how much does latent prediction actually improve data efficiency? Is there a benefit to stacking such methods into multi-scale hierarchies? We answer both using as data a tractable probabilistic context-free grammar that captures the compositional structure of natural language and images. Such a grammar generates strings of visible tokens by recursively applying production rules along a tree of hidden symbols of depth LL. For such data, supervised or token-level SSL require a number of samples \emph{exponential} in LL to recover the latent tree; we prove that latent prediction achieves this with a number of samples \emph{constant} in LL, up to logarithmic factors. We confirm this bound with (i) a hierarchical clustering algorithm, (ii) an end-to-end neural network whose predictor-clusterer modules predict their own latents at each level via gradient descent, and (iii) the first sample-complexity analysis of data2vec, which we show implicitly performs hierarchical latent prediction. This suggests that explicit stacking such as H-JEPA is largely redundant.

Keywords

Cite

@article{arxiv.2605.27734,
  title  = {Learn from your own latents and not from tokens: A sample-complexity theory},
  author = {Daniel J. Korchinski and Alessandro Favero and Matthieu Wyart},
  journal= {arXiv preprint arXiv:2605.27734},
  year   = {2026}
}

Comments

10 pages, 5 figures in main. 28 pages, 14 figures, 1 table in all