English
Related papers

Related papers: Complex Dynamics in Simple Neural Networks: Unders…

200 papers

In gradient descent, changing how we parametrize the model can lead to drastically different optimization trajectories, giving rise to a surprising range of meaningful inductive biases: identifying sparse classifiers or reconstructing…

Machine Learning · Statistics 2021-11-24 Anna Kerekes , Anna Mészáros , Ferenc Huszár

Our work focuses on stochastic gradient methods for optimizing a smooth non-convex loss function with a non-smooth non-convex regularizer. Research on this class of problem is quite limited, and until recently no non-asymptotic convergence…

Optimization and Control · Mathematics 2019-05-15 Michael R. Metel , Akiko Takeda

It seems that in the current age, computers, computation, and data have an increasingly important role to play in scientific research and discovery. This is reflected in part by the rise of machine learning and artificial intelligence,…

Machine Learning · Computer Science 2024-05-15 Ronan Keane

When equipped with efficient optimization algorithms, the over-parameterized neural networks have demonstrated high level of performance even though the loss function is non-convex and non-smooth. While many works have been focusing on…

Machine Learning · Computer Science 2021-03-11 Zhiqi Bu , Shiyun Xu , Kan Chen

We consider the minimization of non-convex quadratic forms regularized by a cubic term, which exhibit multiple saddle points and poor local minima. Nonetheless, we prove that, under mild assumptions, gradient descent approximates the…

Optimization and Control · Mathematics 2022-08-31 Yair Carmon , John C. Duchi

We study the dynamics of a droplet moving on an inclined rough surface in the absence of inertial and viscous stress effects. In this case, the dynamics of the droplet is a purely geometric motion in terms of the wetting domain and the…

Numerical Analysis · Mathematics 2022-11-08 Yuan Gao , Jian-Guo Liu

We present a theoretical and empirical study of the gradient dynamics of overparameterized shallow ReLU networks with one-dimensional input, solving least-squares interpolation. We show that the gradient dynamics of such networks are…

Machine Learning · Computer Science 2019-06-20 Francis Williams , Matthew Trager , Claudio Silva , Daniele Panozzo , Denis Zorin , Joan Bruna

Small-scale turbulence can be comprehensively described in terms of velocity gradients, which makes them an appealing starting point for low-dimensional modeling. Typical models consist of stochastic equations based on closures for…

Fluid Dynamics · Physics 2024-03-01 Maurizio Carbone , Vincent J. Peterhans , Alexander S. Ecker , Michael Wilczek

Perturbative calculations of gradient flow observables are technically challenging. Current results are limited to a few quantities and, in general, to low perturbative orders. Numerical stochastic perturbation theory is a potentially…

High Energy Physics - Lattice · Physics 2016-12-16 Mattia Dalla Brida , Martin Lüscher

Turbulent flows consist of a wide range of interacting scales. Since the scale range increases as some power of the flow Reynolds number, a faithful simulation of the entire scale range is prohibitively expensive at high Reynolds numbers.…

Fluid Dynamics · Physics 2023-07-24 Dhawal Buaria , Katepalli R. Sreenivasan

Motivated by a constrained minimization problem, it is studied the gradient flows with respect to Hessian Riemannian metrics induced by convex functions of Legendre type. The first result characterizes Hessian Riemannian structures on…

Optimization and Control · Mathematics 2018-11-27 Felipe Alvarez , Jérôme Bolte , Olivier Brahic

Phase retrieval consists in the recovery of a complex-valued signal from intensity-only measurements. As it pervades a broad variety of applications, many researchers have striven to develop phase-retrieval algorithms. Classical approaches…

Many iterative procedures in stochastic optimization exhibit a transient phase followed by a stationary phase. During the transient phase the procedure converges towards a region of interest, and during the stationary phase the procedure…

Machine Learning · Statistics 2018-02-26 Jerry Chee , Panos Toulis

We study the problem of estimating low-rank matrices from linear measurements (a.k.a., matrix sensing) through nonconvex optimization. We propose an efficient stochastic variance reduced gradient descent algorithm to solve a nonconvex…

Machine Learning · Statistics 2017-01-17 Xiao Zhang , Lingxiao Wang , Quanquan Gu

Gradient methods are widely used in optimization problems. In practice, while the smoothness parameter can be estimated utilizing techniques such as backtracking, estimating the strong convexity parameter remains a challenge; moreover, even…

Optimization and Control · Mathematics 2026-02-17 Xiaozhe Hu , Sara Pollock , Zhongqin Xue , Yunrong Zhu

A simple model of the driven motion of interacting particles in a two dimensional random medium is analyzed, focusing on the critical behavior near to the threshold that separates a static phase from a flowing phase with a steady-state…

Statistical Mechanics · Physics 2008-02-03 Joe Watson , Daniel S. Fisher

This paper considers the analysis of continuous time gradient-based optimization algorithms through the lens of nonlinear contraction theory. It demonstrates that in the case of a time-invariant objective, most elementary results on…

Optimization and Control · Mathematics 2022-12-23 Patrick M. Wensing , Jean-Jacques E. Slotine

The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-scale data. In this context, momentum-based stochastic optimisation algorithms have become…

Optimization and Control · Mathematics 2024-11-06 Kexin Jin , Jonas Latz , Chenguang Liu , Alessandro Scagliotti

In this paper, we suggest a new framework for analyzing primal subgradient methods for nonsmooth convex optimization problems. We show that the classical step-size rules, based on normalization of subgradient, or on the knowledge of optimal…

Optimization and Control · Mathematics 2023-11-27 Yurii Nesterov

The Discrete Particle Method (DPM) is used to model granular flows down an inclined chute. We observe three major regimes: static piles, steady uniform flows and accelerating flows. For flows over a smooth base, other (quasi-steady) regimes…

Soft Condensed Matter · Physics 2011-08-26 Thomas Weinhart , Anthony Thornton , Stefan Luding , Onno Bokhove