Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training
Abstract
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items , the final classification layer dominates memory, requiring logits and gradients to materialize for a batch of examples. Sampled softmax reduces this cost by restricting the objective to only candidate negative items, resulting in an memory. However, for a fixed budget , it remains unclear whether one should prioritize larger batches or the inclusion of more negative items. We address this question by analyzing sampled-softmax training under a fixed memory constraint. Under standard smoothness and variance assumptions, our theoretical evidence suggests that the fastest convergence arises from an allocation. So, an actionable rule is to include as many objects as possible given computational constraints. Our theory is supported by controlled synthetic and synthetic and four real sequential recommendation benchmarks, including MovieLens-20M. The suggested configuration achieve faster convergence and better final recommendation quality than imbalanced alternatives within the same memory constraint. These findings provide a theoretical and empirical foundation for configuring memory during the training of recommender systems. Code, reproducibility materials, and all scripts for generating figures are available at https://anonymous.4open.science/r/LimitedMemoryRule-BBFB
Keywords
Cite
@article{arxiv.2608.11061,
title = {Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training},
author = {Artyom Sabitov and Daniil Volkov and Alexey Zaytsev},
journal= {arXiv preprint arXiv:2608.11061},
year = {2026}
}