English
Related papers

Related papers: First-Passage Prediction of Grokking Delay: ACalib…

200 papers

Standard optimization theories struggle to explain grokking, where generalization occurs long after training convergence. While geometric studies attribute this to slow drift, they often overlook the interaction between the optimizer's…

Machine Learning · Computer Science 2026-03-17 Pratyush Acharya , Habish Dhakal

Adaptive optimizers with decoupled weight decay, such as AdamW, are the de facto standard for pre-training large transformer-based generative models. Yet the quadratic nature of the $\ell_2$ penalty embedded in weight decay drives all…

Machine Learning · Computer Science 2025-11-19 Fu-Ming Guo , Yingfang Fan

Grokking -- the abrupt transition from memorization to generalization after prolonged training -- has been linked to confinement on low-dimensional execution manifolds in modular arithmetic. Whether this mechanism extends beyond arithmetic…

Machine Learning · Computer Science 2026-04-06 Yongzhong Xu

Grokking, a delayed generalization in neural networks after perfect training performance, has been observed in Transformers and MLPs, but the components driving it remain underexplored. We show that embeddings are central to grokking:…

Machine Learning · Computer Science 2025-05-22 H. V. AlquBoj , Hilal AlQuabeh , Velibor Bojkovic , Munachiso Nwadike , Kentaro Inui

Grokking -- the delayed transition from memorization to generalization in small algorithmic tasks -- remains poorly understood. We present a geometric analysis of optimization dynamics in transformers trained on modular arithmetic. PCA of…

Machine Learning · Computer Science 2026-04-06 Yongzhong Xu

We explore first-passage phenomenology for biased active processes with a renewal-type structure, focusing in particular on paradigmatic run-and-tumble models in both discrete and continuous state spaces. In general, we show there is no…

Statistical Mechanics · Physics 2025-12-09 Yonathan Sarmiento , Benjamin Walter , Debraj Das , Samvit Mahapatra , Édgar Roldán , Rosemary J. Harris

\emph{Memorization} in neural networks lacks a precise operational definition and is often inferred from the grokking regime, where training accuracy saturates while test accuracy remains very low. We identify a previously unreported third…

Machine Learning · Computer Science 2026-02-04 Hari K Prakash , Charles H Martin

Variational approximation, such as mean-field (MF) and tree-reweighted (TRW), provide a computationally efficient approximation of the log-partition function for a generic graphical model. TRW provably provides an upper bound, but the…

Data Structures and Algorithms · Computer Science 2021-08-23 Romain Cosson , Devavrat Shah

As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-understood. This paper establishes the convergence rate…

Machine Learning · Computer Science 2025-10-06 Huan Li , Yiming Dong , Zhouchen Lin

We use measurements from the Planck satellite mission and galaxy redshift surveys over the last decade to test three of the basic assumptions of the standard model of cosmology, $\Lambda$CDM: the spatial curvature of the universe, the…

Cosmology and Nongalactic Astrophysics · Physics 2016-01-20 Shadab Alam , Shirley Ho , Alessandra Silvestri

Grokking refers to delayed generalization in which the increase in test accuracy of a neural network occurs appreciably after the improvement in training accuracy This paper introduces several practical metrics including variance under…

Machine Learning · Computer Science 2025-07-17 Ahmed Salah , David Yevick

We compare the predictions of hybrid inflationary models that produce both adiabatic fluctuations and topological defects to first year WMAP results. We use a Markov Chain Monte Carlo method to constrain the contribution of cosmic strings…

Astrophysics · Physics 2007-05-23 Aurélien A. Fraisse

We study the well-known grokking phenomena in neural networks (NNs) using a 3-layer MLP trained on 1 k-sample subset of MNIST, with and without weight decay, and discover a novel third phase -- \emph{anti-grokking} -- that occurs very late…

Machine Learning · Computer Science 2025-06-06 Hari K. Prakash , Charles H. Martin

We use the gravitational wave (GW) events GW170817 and GW190521, together with their proposed electromagnetic counterparts, to constrain cosmological parameters and theories of gravity beyond General Relativity (GR). In particular we…

General Relativity and Quantum Cosmology · Physics 2021-03-03 S. Mastrogiovanni , L. Haegel , C. Karathanasis , I. Magana-Hernandez , D. A. Steer

Grokking describes a delayed generalization phenomenon in which a neural network achieves perfect training accuracy long before validation accuracy improves, followed by an abrupt transition to strong generalization. Existing detection…

Machine Learning · Computer Science 2026-04-24 Shreel Golwala

Grokking, the phenomenon of delayed generalization, is often attributed to the depth and compositional structure of deep neural networks. We study grokking in one of the simplest possible settings: the learning of a linear model with…

Machine Learning · Computer Science 2026-02-10 Nataraj Das , Atreya Vedantam , Chandrashekar Lakshminarayanan

Large language models (LLMs) are notoriously memory-intensive during training, particularly with the popular AdamW optimizer. This memory burden necessitates using more or higher-end GPUs or reducing batch sizes, limiting training…

Machine Learning · Computer Science 2025-02-18 Hanqing Zhu , Zhenyu Zhang , Wenyan Cong , Xi Liu , Sem Park , Vikas Chandra , Bo Long , David Z. Pan , Zhangyang Wang , Jinwon Lee

Empirical scaling laws prescribe how to allocate parameters, data, and compute, while maximal-update parameterization ($\mu$P) enables learning-rate transfer across widths by equalizing early-time update magnitudes. However, in modern…

Machine Learning · Computer Science 2025-10-20 Zhiyuan Fan , Yifeng Liu , Qingyue Zhao , Angela Yuan , Quanquan Gu

While world models learn compact representations of complex environments, they lack a physics-grounded metric to assess the structural fidelity of their latent spaces. We identify the wavelet scaling exponent $\alpha$ as a critical…

Quantum Physics · Physics 2026-05-13 Chon-Fai Kam , Xavier Cadet , Miloud Bessafi , Frederic Cadet

We investigate the work fluctuations in an overdamped non-equilibrium process that is stopped at a stochastic time. The latter is characterized by a first passage event that marks the completion of the non-equilibrium process. In…

Statistical Mechanics · Physics 2024-03-20 Iago N Mamede , Prashant Singh , Arnab Pal , Carlos E. Fiore , Karel Proesmans
‹ Prev 1 2 3 10 Next ›