Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization
Abstract
Weight decay is widely used as a regularizer in large language models, yet its precise role in shaping Transformer loss landscapes remains theoretically underexplored. This paper provides the first rigorous functional-analytic characterization of the standard Transformer objective--cross-entropy loss with regularization--by proving it satisfies Villani's criteria for coercive energy functions. Specifically, we show that the regularized loss is infinitely differentiable, grows at least quadratically, has Gaussian-integrable tails, and satisfies the differential growth condition as for all . From this structure, we derive explicit log-Sobolev and Poincar\'e constants , linking the regularization strength and model dimension to finite-time convergence guarantees for noisy stochastic gradient descent and PAC-Bayesian generalization bounds that tighten with increasing . To validate our theory, we introduce a scalable Villani diagnostic and estimate it efficiently using Hutchinson trace probes in models with over 100M parameters. Experiments on GPT-Neo-125M across Penn Treebank and WikiText-103 confirm the predicted quadratic growth of , spectral inflation of the Hessian, and exponential convergence behavior consistent with our log-Sobolev analysis. These results demonstrate that weight decay not only improves generalization empirically but also establishes the mathematical conditions required for fast Langevin mixing and theoretically grounded curvature-aware optimization in deep learning.
Keywords
Cite
@article{arxiv.2605.06599,
title = {Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization},
author = {Abhijit Das and Sayantan Dutta},
journal= {arXiv preprint arXiv:2605.06599},
year = {2026}
}
Comments
17 pages, 10 figures