Refresh-Scaling the Memory of Balanced Adam
Abstract
Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, , reducing the optimizer to a single remaining parameter. However, how this parameter should be set remains poorly understood. We argue that, in balanced Adam, should not be treated as a dimensionless constant: it defines a statistical memory horizon . In terms of the effective learning horizon , estimated from the validation trajectory, we study the refresh count , which measures how many times Adam renews its internal statistics during the useful phase of training. Across 11 vision and language experiments, we find that choosing so that selects different values depending on the training scale, yet improves robustness over the best fixed-beta baseline. Compared with the strongest fixed choice , the refresh rule improves worst-case robustness, reducing the maximum relative gap in validation loss by 33.4\%, while bringing all 11 runs within 1\% of their validation oracle. These results suggest that the remaining hyperparameter of balanced Adam is more naturally viewed as a memory-scale variable than as a fixed constant. This provides a simple budget-aware perspective on optimizer scaling and opens a path toward treating Adam's momentum as part of the learning dynamics rather than as a static default.
Keywords
Cite
@article{arxiv.2605.10119,
title = {Refresh-Scaling the Memory of Balanced Adam},
author = {Alberto Fernández-Hernández and Cristian Pérez-Corral and Jose I. Mestre and Manuel F. Dolz and Enrique S. Quintana-Ortí},
journal= {arXiv preprint arXiv:2605.10119},
year = {2026}
}