Related papers: Global Convergence of Second-order Dynamics in Two…
We present a unified convergence analysis for first order convex optimization methods using the concept of strong Lyapunov conditions. Combining this with suitable time scaling factors, we are able to handle both convex and strong convex…
We introduce new multilevel methods for solving large-scale unconstrained optimization problems. Specifically, the philosophy of multilevel methods is applied to Newton-type methods that regularize the Newton sub-problem using second order…
In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution at a linear convergence rate for the quadratic loss…
In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in minimizing nonsmooth nonconvex functions. We first investigate the global stability of a general framework for stochastic subgradient methods,…
We study the overparametrization bounds required for the global convergence of stochastic gradient descent algorithm for a class of one hidden layer feed-forward neural networks, considering most of the activation functions used in…
We give a simple proof for the global convergence of gradient descent in training deep ReLU networks with the standard square loss, and show some of its improvements over the state-of-the-art. In particular, while prior works require all…
Classical global convergence results for first-order methods rely on uniform smoothness and the \L{}ojasiewicz inequality. Motivated by properties of objective functions that arise in machine learning, we propose a non-uniform refinement of…
We prove that the Gini coefficient of economic inequality is a Lyapunov functional for a class of nonlinear, nonlocal integro-differential equations arising at the intersection of mathematics, economics, and statistical physics. Next, a…
We consider a one dimensional transport model with nonlocal velocity given by the Hilbert transform and develop a global well-posedness theory of probability measure solutions. Both the viscous and non-viscous cases are analyzed. Both in…
We consider a one-dimensional kinetic model of granular media in the case where the interaction potential is quadratic. Taking advan- tage of a simple first integral, we can use a reformulation (equivalent to the initial kinetic model for…
We show that the Nernst-Planck-Euler system, which models ionic electrodiffusion in fluids, has global strong solutions for arbitrarily large data in the two dimensional bounded domains. The assumption on species is either there are two…
We propose an unconstrained optimization method based on the well-known primal-dual hybrid gradient (PDHG) algorithm. We first formulate the optimality condition of the unconstrained optimization problem as a saddle point problem. We then…
We present a novel notion of $\lambda$-monotonicity for an $n$-species system of partial differential equations governed by mass-preserving flow dynamics, extending monotonicity in Banach spaces to the Wasserstein-2 metric space. We show…
Deep learning models are often successfully trained using gradient descent, despite the worst case hardness of the underlying non-convex optimization problem. The key question is then under what conditions can one prove that optimization…
There are much recent interests in solving noncovnex min-max optimization problems due to its broad applications in many areas including machine learning, networked resource allocations, and distributed optimization. Perhaps, the most…
Over the past fifteen years, the theory of Wasserstein gradient flows of convex (or, more generally, semiconvex) energies has led to advances in several areas of partial differential equations and analysis. In this work, we extend the…
The symmetric low-rank matrix factorization serves as a building block in many learning tasks, including matrix recovery and training of neural networks. However, despite a flurry of recent research, the dynamics of its training via…
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…
Minimizing loss functions is central to machine-learning training. Although first-order methods dominate practical applications, higher-order techniques such as Newton's method can deliver greater accuracy and faster convergence, yet are…
We introduce a framework for Newton's flows in probability space with information metrics, named information Newton's flows. Here two information metrics are considered, including both the Fisher-Rao metric and the Wasserstein-2 metric. A…