Related papers: Gradient descent in higher codimension
We provide a covariant and gauge-invariant approach to the question of how a first order pressure can be incorporated self-consistently in a cosmological scenario. The approximation is relevant, in the linear regime, to weakly…
An experimental arrangement and a set of experiments are developed to generate empirical evidence of the effect of noise on a rotating, macro-scale cantilever structure. The experiment is a controlled representation of a rotating machinery…
We have analyzed the effects of the addition of external noise to non-dynamical systems displaying intrinsic noise, and established general conditions under which stochastic resonance appears. The criterion we have found may be applied to a…
Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking. These phenomena appear across architectures -- in…
In this study, we explore the effects of including noise predictors and noise observations when fitting linear regression models. We present empirical and theoretical results that show that double descent occurs in both cases, albeit with…
Fractional derivatives are a well-studied generalization of integer order derivatives. Naturally, for optimization, it is of interest to understand the convergence properties of gradient descent using fractional derivatives. Convergence…
Convolutional neural networks are widely used in imaging and image recognition. Learning such networks from training data leads to the minimization of a non-convex function. This makes the analysis of standard optimization methods such as…
Finding a local minimum or maximum of a function is often achieved through the gradient-descent optimization method. For a function in dimension d, the gradient requires to compute at each step d partial derivatives. This method is for…
Noisy dynamical models are employed to describe a wide range of phenomena. Since exact modeling of these phenomena requires access to their microscopic dynamics, whose time scales are typically much shorter than the observable time scales,…
Stochastic Gradient Descent (SGD) is the main approach to optimizing neural networks. Several generalization properties of deep networks, such as convergence to a flatter minima, are believed to arise from SGD. This article explores the…
It is important to understand how the popular regularization method dropout helps the neural network training find a good generalization solution. In this work, we show that the training with dropout finds the neural network with a flatter…
This paper establishes risk convergence and asymptotic weight matrix alignment --- a form of implicit regularization --- of gradient flow and gradient descent when applied to deep linear networks on linearly separable data. In more detail,…
The dynamical stability of the iterates during training plays a key role in determining the minima obtained by optimization algorithms. For example, stable solutions of gradient descent (GD) correspond to flat minima, which have been…
Understanding when the noise in stochastic gradient descent (SGD) affects generalization of deep neural networks remains a challenge, complicated by the fact that networks can operate in distinct training regimes. Here we study how the…
As children grow older, they develop an intuitive understanding of the physical processes around them. They move along developmental trajectories, which have been mapped out extensively in previous empirical research. We investigate how…
Gradient descent can be surprisingly good at optimizing deep neural networks without overfitting and without explicit regularization. We find that the discrete steps of gradient descent implicitly regularize models by penalizing gradient…
Sliding motion is evolution on a switching manifold of a discontinuous, piecewise-smooth system of ordinary differential equations. In this paper we quantitatively study the effects of small-amplitude, additive, white Gaussian noise on…
The dynamics of droplets on substrates has a strong impact on microfluidic systems ranging from commercially available lab-on-chip systems to state of the art developments in open microfluidics. Coalescence of micro and nano droplets on a…
We explicitly construct parameter transformations between gradient flows in metric spaces, called curves of maximal slope, having different exponents when the associated function satisfies a suitable convexity condition. These…
We formulate and study a general family of (continuous-time) stochastic dynamics for accelerated first-order minimization of smooth convex functions. Building on an averaging formulation of accelerated mirror descent, we propose a…