Related papers: Noether: The More Things Change, the More Stay the…
A symmetry is a `change without change'. As simple as it sounds, this concept is the fundamental cornerstone that unifies all branches of theoretical physics. Virtually all physical laws -- ranging from classical mechanics and…
In this thesis, we study the one parameter point transformations which leave invariant the differential equations. In particular we study the Lie and the Noether point symmetries of second order differential equations. We establish a new…
When training deep neural networks with gradient descent, sharpness often increases -- a phenomenon known as progressive sharpening -- before saturating at the edge of stability. Although commonly observed in practice, the underlying…
We propose an empirical approach centered on the spectral dynamics of weights -- the behavior of singular values and vectors during optimization -- to unify and clarify several phenomena in deep learning. We identify a consistent bias in…
We sketch the main features of the Noether Symmetry Approach, a method to reduce and solve dynamics of physical systems by selecting Noether symmetries, which correspond to conserved quantities. Specifically, we take into account the…
Understanding complex systems with their reduced model is one of the central roles in scientific activities. Although physics has greatly been developed with the physical insights of physicists, it is sometimes challenging to build a…
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…
When discussing consequences of symmetries of dynamical systems based on Noether's first theorem, most standard textbooks on classical or quantum mechanics present a conclusion stating that a global continuous Lie symmetry implies the…
Neural Persistence is a prominent measure for quantifying neural network complexity, proposed in the emerging field of topological data analysis in deep learning. In this work, however, we find both theoretically and empirically that the…
When an online learning algorithm is used to estimate the unknown parameters of a model, the signals interacting with the parameter estimates should not decay too quickly for the optimal values to be discovered correctly. This requirement…
We consider Convolutional Neural Networks (CNNs) with 2D structured features that are symmetric in the spatial dimensions. Such networks arise in modeling pairwise relationships for a sequential recommendation problem, as well as secondary…
This paper discusses the interplay of symmetries and stability in the analysis and control of nonlinear dynamical systems and networks. Specifically, it combines standard results on symmetries and equivariance with recent convergence…
Dependent symmetries, symmetries that depend on the situation of the subsystem in a larger closed system, are explored by looking at simple examples. This is a new kind of symmetry in the open quantum dynamics of a subsystem Each symmetry…
Neural network optimization remains one of the most consequential yet poorly understood challenges in modern AI research, where improvements in training algorithms can lead to enhanced feature learning in foundation models,…
Neural networks appear to have mysterious generalization properties when using parameter counting as a proxy for complexity. Indeed, neural networks often have many more parameters than there are data points, yet still provide good…
The first part of this paper develops a geometric setting for differential-difference equations that resolves an open question about the extent to which continuous symmetries can depend on discrete independent variables. For general…
Symmetries have been leveraged to improve the generalization of neural networks through different mechanisms from data augmentation to equivariant architectures. However, despite their potential, their integration into neural solvers for…
A main puzzle of deep networks revolves around the absence of overfitting despite large overparametrization and despite the large capacity demonstrated by zero training error on randomly labeled data. In this note, we show that the dynamics…
Expressiveness and generalization of deep models was recently addressed via the connection between neural networks (NNs) and kernel learning, where first-order dynamics of NN during a gradient-descent (GD) optimization were related to…
Recent work on mode connectivity in the loss landscape of deep neural networks has demonstrated that the locus of (sub-)optimal weight vectors lies on continuous paths. In this work, we train a neural network that serves as a hypernetwork,…