Related papers: Insights on Muon from Simple Quadratics
The ever-growing scale of deep learning models and training data underscores the critical importance of efficient optimization methods. While preconditioned gradient methods such as Adam and AdamW are the de facto optimizers for training…
We present new results for the light-quark connected part of the leading order hadronic-vacuum-polarization (HVP) contribution to the muon anomalous magnetic moment, using $2+1+1$ staggered fermions. We have collected more statistics on…
The hadronic vacuum polarization contribution to the anomalous magnetic moment of the muon is parametrized by using the quark-resonance model formulated in \cite{QR}. In this context a recent prediction obtained within the ENJL model…
This paper investigates the impact of different optimizers on the grokking phenomenon, where models exhibit delayed generalization. We conducted experiments across seven numerical tasks (primarily modular arithmetic) using a modern…
With the recent interest in the measured and standard model values of the muon anomalous magnetic moment, a_mu, some confusion has arisen concerning our knowledge of the hadronic contribution to a_mu. In the dispersion integral approach to…
Human perception of the empirical world involves recognizing the diverse appearances, or 'modalities', of underlying objects. Despite the longstanding consideration of this perspective in philosophy and cognitive science, the study of…
The anomalous magnetic moment of the muon displays a $4.2\sigma$ tension with the Standard-Model prediction, if $e^+e^-\to \text{hadrons}$ data are used for hadronic vacuum polarization. In these proceedings we review possible explanations…
We present a quenched lattice calculation of the lowest order (alpha^2) hadronic contribution to the anomalous magnetic moment of the muon which arises from the hadronic vacuum polarization. A general method is presented for computing…
The hadronic contributions to the muon anomalous magnetic moment and to the shift of the electromagnetic fine structure constant at the scale of Z boson mass are evaluated within dispersively improved perturbation theory (DPT). The latter…
Statistical inferences for quadratic functionals of linear regression parameter have found wide applications including signal detection, global testing, inferences of error variance and fraction of variance explained. Classical theory based…
Successive quadratic approximations, or second-order proximal methods, are useful for minimizing functions that are a sum of a smooth part and a convex, possibly nonsmooth part that promotes regularization. Most analyses of iteration…
We investigate the use of effective Lagrangians to describe the effects on high-precision observables of physics beyond the Standard Model. Using the anomalous magnetic moment of the muon as an example, we detail the use of effective…
We consider a minimal realization of a rational matrix functions. We perturb the polynomial part and one of the constant matrices from the realization part. We derive explicit computable expressions of backward errors of approximate…
We consider Proximal Newton methods with an inexact computation of update steps. To this end, we introduce two inexactness criteria which characterize sufficient accuracy of these update step and with the aid of these investigate global…
Foundational Machine Learning Potentials can resolve the accuracy and transferability limitations of classical force fields. They enable microscopic insights into material behavior through Molecular Dynamics simulations, which can crucially…
Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlooked. Recently, the optimizer Muon \citep{jordanmuon}, which explicitly exploits this…
Large models recently are widely applied in artificial intelligence, so efficient training of large models has received widespread attention. More recently, a useful Muon optimizer is specifically designed for matrix-structured parameters…
Given a matrix $A$, a linear feasibility problem (of which linear classification is a special case) aims to find a solution to a primal problem $w: A^Tw > \textbf{0}$ or a certificate for the dual problem which is a probability distribution…
Functional factor analysis is an important dimension reduction method for functional and longitudinal data. Factor loadings give insight into patterns of variability of the observations, while latent factors provide a low-dimensional…
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iterations as its core operation. Existing Muon variants apply a uniform NS schedule to all…