Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Machine Learning
2026-07-08 v1
Abstract
In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer depth. We show that even in this minimal model, the optimal depth-wise learning rate scaling depends on data, whereas data-agnostic scaling rules fail to transfer across depths. Under the data-dependent optimal scaling, the learning dynamics is independent of data and weakly dependent on depth, resulting in a constant linear convergence rate across all depths including infinity. We further show similar data-dependent effects in deep scalar linear networks with residual connections.
Cite
@article{arxiv.2607.07884,
title = {Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks},
author = {Yedi Zhang and Peter E. Latham and Leena Chennuru Vankadara and Andrew Saxe},
journal= {arXiv preprint arXiv:2607.07884},
year = {2026}
}