English
Related papers

Related papers: Gradient descent in some simple settings

200 papers

Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking. These phenomena appear across architectures -- in…

Machine Learning · Computer Science 2026-01-01 Alan Oursland

A tutorial review is given of some developments and applications of stochastic processes from the point of view of the practicioner physicist. The index is the following: 1.- Introduction 2.- Stochastic Processes 3.- Transient Stochastic…

Condensed Matter · Physics 2007-05-23 Maxi San Miguel , Raul Toral

The starting assumptions to study the convergence and complexity of gradient-type methods may be the smoothness (also called Lipschitz continuity of gradient) and the strong convexity. In this note, we revisit these two basic properties…

Optimization and Control · Mathematics 2021-11-01 Lu Zhang , Jiani Wang , Hui Zhang

Differentially private stochastic gradient descent (DP-SGD) is known to have poorer training and test performance on large neural networks, compared to ordinary stochastic gradient descent (SGD). In this paper, we perform a detailed study…

Machine Learning · Computer Science 2023-11-14 Lauren Watson , Eric Gan , Mohan Dantam , Baharan Mirzasoleiman , Rik Sarkar

Many tasks in machine learning and signal processing can be solved by minimizing a convex function of a measure. This includes sparse spikes deconvolution or training a neural network with a single hidden layer. For these problems, we study…

Optimization and Control · Mathematics 2018-10-30 Lenaic Chizat , Francis Bach

We present a new class of gradient-type optimization methods that extends vanilla gradient descent, mirror descent, Riemannian gradient descent, and natural gradient descent. Our approach involves constructing a surrogate for the objective…

Optimization and Control · Mathematics 2023-06-13 Flavien Léger , Pierre-Cyril Aubin-Frankowski

We investigate the effect of time-dependent noise on the shape of a morphogen gradient in a developing embryo. Perturbation theory is used to calculate the deviations from deterministic behavior in a simple reaction-diffusion model of…

Tissues and Organs · Quantitative Biology 2007-05-23 Jeremy L. England , John Cardy

We consider a slow passage through a point of loss of stability. If the passage is sufficiently slow, the dynamics are controlled by additive random disturbances, even if they are extremely small. We derive expressions for the `exit value'…

adap-org · Physics 2008-02-03 G. D. Lythe

The paper contributes to strengthening the relation between machine learning and the theory of differential equations. In this context, the inverse problem of fitting the parameters, and the initial condition of a differential equation to…

Machine Learning · Computer Science 2022-06-22 Imre Fekete , András Molnár , Péter L. Simon

Although the theoretical behavior of one-dimensional random walks in random environments is well understood, the numerical evaluation of various characteristics of such processes has received relatively little attention. This paper develops…

Probability · Mathematics 2014-06-16 Werner R. W. Scheinhardt , Dirk P. Kroese

How to find flat minima? We propose running normalized gradient descent, usually reserved for nonsmooth optimization, with sufficiently slowly diminishing step sizes. This induces implicit regularization towards flat minima if an…

Optimization and Control · Mathematics 2026-02-10 Cédric Josz

In machine learning, stochastic gradient descent (SGD) is widely deployed to train models using highly non-convex objectives with equally complex noise models. Unfortunately, SGD theory often makes restrictive assumptions that fail to…

Machine Learning · Computer Science 2022-10-11 Vivak Patel , Shushu Zhang , Bowen Tian

Additive noise in Partial Differential equations, in particular those of fluid mechanics, has relatively natural motivations. The aim of this work is showing that suitable multiscale arguments lead rigorously, from a model of fluid with…

Probability · Mathematics 2022-05-12 Franco Flandoli , Umberto Pappalettera

Noise is often considered to be a nuisance. Here we argue that it can be a useful probe of fluctuating two level systems in glasses. It can be used to: (1) shed light on whether the fluctuations are correlated or independent events; (2)…

Disordered Systems and Neural Networks · Physics 2009-11-10 Clare C. Yu

In this contribution we present an exploratory study of several novel methods for numerical stochastic perturbation theory. For the investigation we consider observables defined through the gradient flow in the simple {\phi}^4 theory.

High Energy Physics - Lattice · Physics 2015-12-29 Mattia Dalla Brida , Marco Garofalo , Anthony D. Kennedy

In the first part of this paper, we consider a family of continuous-time dynamical systems coupled with diffusion-transmutation processes. Under certain conditions, such randomly perturbed dynamical systems can be interpreted as an averaged…

Optimization and Control · Mathematics 2024-08-21 Getachew K. Befekadu

How much can you say about the gradient of a neural network without computing a loss or knowing the label? This may sound like a strange question: surely the answer is "very little." However, in this paper, we show that gradients are more…

Microcanonical gradient descent is a sampling procedure for energy-based models allowing for efficient sampling of distributions in high dimension. It works by transporting samples from a high-entropy distribution, such as Gaussian white…

Machine Learning · Statistics 2024-05-28 Marcus Häggbom , Morten Karlsmark , Joakim Andén

Gradient dynamics play a central role in determining the stability and generalization of deep neural networks. In this work, we provide an empirical analysis of how variance and standard deviation of gradients evolve during training,…

Machine Learning · Computer Science 2025-09-09 Vincent-Daniel Yun

We describe an alternative learning method for neural networks, which we call Blind Descent. By design, Blind Descent does not face problems like exploding or vanishing gradients. In Blind Descent, gradients are not used to guide the…

Machine Learning · Computer Science 2020-08-27 Akshat Gupta , Prasad N R
‹ Prev 1 4 5 6 7 8 10 Next ›