English

Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs

Machine Learning 2024-09-06 v1 Machine Learning

Abstract

We study the impact of the batch size nbn_b on the iteration time TT of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target functions of isotropic covariates. We characterize the optimal batch size minimizing the iteration time as a function of the hardness of the target, as characterized by the information exponents. We show that performing gradient updates with large batches nbd2n_b \lesssim d^{\frac{\ell}{2}} minimizes the training time without changing the total sample complexity, where \ell is the information exponent of the target to be learned \citep{arous2021online} and dd is the input dimension. However, larger batch sizes than nbd2n_b \gg d^{\frac{\ell}{2}} are detrimental for improving the time complexity of SGD. We provably overcome this fundamental limitation via a different training protocol, \textit{Correlation loss SGD}, which suppresses the auto-correlation terms in the loss function. We show that one can track the training progress by a system of low-dimensional ordinary differential equations (ODEs). Finally, we validate our theoretical results with numerical experiments.

Keywords

Cite

@article{arxiv.2406.02157,
  title  = {Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs},
  author = {Luca Arnaboldi and Yatin Dandi and Florent Krzakala and Bruno Loureiro and Luca Pesce and Ludovic Stephan},
  journal= {arXiv preprint arXiv:2406.02157},
  year   = {2024}
}