English
Related papers

Related papers: Gradient descent in higher codimension

200 papers

We consider in this work small random perturbations (of multiplicative noise type) of the gradient flow. We prove that under mild conditions, when the potential function is a Morse function with additional strong saddle condition, the…

Probability · Mathematics 2020-04-29 Jiaojiao Yang , Wenqing Hu , Chris Junchi Li

Differentially private (DP) release of multidimensional statistics typically considers an aggregate sensitivity, e.g. the vector norm of a high-dimensional vector. However, different dimensions of that vector might have widely different…

Machine Learning · Statistics 2022-10-31 Joonas Jälkö , Lukas Prediger , Antti Honkela , Samuel Kaski

How much can you say about the gradient of a neural network without computing a loss or knowing the label? This may sound like a strange question: surely the answer is "very little." However, in this paper, we show that gradients are more…

In this paper we study a gradient flow approach to the problem of quantization of measures in one dimension. By embedding our problem in $L^2$, we find a continuous version of it that corresponds to the limit as the number of particles…

Analysis of PDEs · Mathematics 2016-01-26 Emanuele Caglioti , François Golse , Mikaela Iacobelli

Low-precision training has become crucial for reducing the computational and memory costs of large-scale deep learning. However, quantizing gradients introduces magnitude shrinkage, which can change how stochastic gradient descent (SGD)…

Machine Learning · Computer Science 2026-01-09 Vincent-Daniel Yun

We show that gradient descent converges to a local minimizer, almost surely with random initialization. This is proved by applying the Stable Manifold Theorem from dynamical systems theory.

Machine Learning · Statistics 2016-03-07 Jason D. Lee , Max Simchowitz , Michael I. Jordan , Benjamin Recht

Applying Differentially Private Stochastic Gradient Descent (DPSGD) to training modern, large-scale neural networks such as transformer-based models is a challenging task, as the magnitude of noise added to the gradients at each iteration…

Machine Learning · Computer Science 2022-07-07 Ryuichi Ito , Seng Pei Liew , Tsubasa Takahashi , Yuya Sasaki , Makoto Onizuka

Neural networks trained with stochastic gradient descent exhibit an inductive bias towards simpler decision boundaries, typically converging to a narrow family of functions, and often fail to capture more complex features. This phenomenon…

Machine Learning · Computer Science 2024-11-08 Rahul Vashisht , P. Krishna Kumar , Harsha Vardhan Govind , Harish G. Ramaswamy

We study the numerical behaviour of a particle method for gradient flows involving linear and nonlinear diffusion. This method relies on the discretisation of the energy via non-overlapping balls centred at the particles. The resulting…

Analysis of PDEs · Mathematics 2016-12-07 J. A. Carrillo , Y. Huang , F. S. Patacchini , G. Wolansky

A discrete system constituted of particles interacting by means of a centroid-based law is numerically investigated. The elements of the system move in the plane, and the range of the interaction can be varied from a more local form…

Soft Condensed Matter · Physics 2020-04-22 A. Battista , L. Rosa , L. Greco , R. dell'Erba

We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a strong component along these directions. Furthermore, we show…

Machine Learning · Computer Science 2018-09-18 Hadi Daneshmand , Jonas Kohler , Aurelien Lucchi , Thomas Hofmann

We study stochastic policy gradient methods from the perspective of control-theoretic limitations. Our main result is that ill-conditioned linear systems in the sense of Doyle inevitably lead to noisy gradient estimates. We also give an…

Optimization and Control · Mathematics 2022-06-15 Ingvar Ziemann , Anastasios Tsiamis , Henrik Sandberg , Nikolai Matni

This paper investigates the behavior of the Min-Sum decoder running on noisy devices. The aim is to evaluate the robustness of the decoder in the presence of computation noise, e.g. due to faulty logic in the processing units, which…

Information Theory · Computer Science 2014-05-27 Christiane L. Kameni Ngassa , Valentin Savin , Elsa Dupraz , David Declercq

We present a new numerical scheme for one dimensional dynamical systems. This is a modification of the discrete gradient method and keeps its advantages, including the stability and the conservation of the energy integral. However, its…

Numerical Analysis · Computer Science 2015-05-13 Jan L. Cieslinski , Boguslaw Ratkiewicz

This work analyzes the convergence of a class of smoothing-based gradient descent methods when applied to optimization problems. In particular, Gaussian smoothing is employed to define a nonlocal gradient that reduces high-frequency noise,…

Optimization and Control · Mathematics 2024-03-27 Andrew Starnes , Anton Dereventsov , Clayton Webster

We study differentiable strongly quasiconvex functions for providing new properties for algorithmic and monotonicity purposes. Furthemore, we provide insights into the decreasing behaviour of strongly quasiconvex functions, applying this…

Optimization and Control · Mathematics 2024-10-07 Felipe Lara , Raúl T. Marcavillaca , Phan T. Vuong

Stochastic gradient descent is an optimisation method that combines classical gradient descent with random subsampling within the target functional. In this work, we introduce the stochastic gradient process as a continuous-time…

Probability · Mathematics 2021-05-11 Jonas Latz

Protecting privacy in learning while maintaining the model performance has become increasingly critical in many applications that involve sensitive data. Private Gradient Descent (PGD) is a commonly used private learning framework, which…

Machine Learning · Computer Science 2022-10-20 Junyuan Hong , Zhangyang Wang , Jiayu Zhou

The coarsening process in a class of driven systems is studied. These systems have previously been shown to exhibit phase separation and slow coarsening in one dimension. We consider generalizations of this class of models to higher…

Statistical Mechanics · Physics 2015-06-24 Y. Kafri , D. Biron , M. R. Evans , D. Mukamel

This paper develops methodology for local sensitivity analysis based on directional derivatives associated with spatial processes. Formal gradient analysis for spatial processes was elaborated in previous papers, focusing on distribution…

Statistics Theory · Mathematics 2015-03-31 Maria A. Terres , Alan E. Gelfand
‹ Prev 1 3 4 5 6 7 10 Next ›