Related papers: Adjacency Criterion For Gradient Flow With Multipl…
In this article we consider an optimization problem where the objective function is evaluated at the fixed-point of a contraction mapping parameterized by a control variable, and optimization takes place over this control variable. Since…
We consider the popular and classical method of alternating projections for finding a point in the intersection of two closed sets. By situating the algorithm in a metric space, equipped only with well-behaved geodesics and angles (in the…
Readability criteria, such as distance or neighborhood preservation, are often used to optimize node-link representations of graphs to enable the comprehension of the underlying data. With few exceptions, graph drawing algorithms typically…
We are interested in existence of gradient flows for shape functionals especially for first Laplacian eigenvalues. We introduce different techniques to prove existence and use different formulations for gradient flows. We apply a…
Bilevel optimization has been developed for many machine learning tasks with large-scale and high-dimensional data. This paper considers a constrained bilevel optimization problem, where the lower-level optimization problem is convex with…
Optimization is at the heart of machine learning, statistics and many applied scientific disciplines. It also has a long history in physics, ranging from the minimal action principle to finding ground states of disordered systems such as…
Let $(M,g)$ be a closed Riemannian manifold, and let $F:M \to \mathbb{R}$ be a smooth function on $M$. We show the following holds generically for the function $F$: for each maximum $p$ of $F$, there exist two minima, denoted by $m_+(p)$…
This paper studies stochastic control problems with the action space taken to be probability measures, with the objective penalised by the relative entropy. We identify suitable metric space on which we construct a gradient flow for the…
We study the training dynamics of shallow neural networks, in a two-timescale regime in which the stepsizes for the inner layer are much smaller than those for the outer layer. In this regime, we prove convergence of the gradient flow to a…
We introduce a site-wise domination criterion for local percolation models, which enables the comparison of one-arm probabilities even in the absence of stochastic domination. The method relies on a local-to-global principle: if, at each…
We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence…
The goal of this paper is to design optimal multilevel solvers for the finite element approximation of second order linear elliptic problems with piecewise constant coefficients on bisection grids. Local multigrid and BPX preconditioners…
We prove the well-posedness of entropy solutions for a wide class of nonlocal transport equations with nonlinear mobility in one spatial dimension. The solution is obtained as the limit of approximations constructed via a deterministic…
This paper applies a discrete adjoint gradient computation method for a multi-class traffic flow model on road networks. Vehicle classes are characterized by their specific velocity functions, which depend on the total traffic density,…
We investigate the test risk of continuous-time stochastic gradient flow dynamics in learning theory. Using a path integral formulation we provide, in the regime of a small learning rate, a general formula for computing the difference…
We study the dynamics of gradient flow for training a multi-head softmax attention model for in-context learning of multi-task linear regression. We establish the global convergence of gradient flow under suitable choices of initialization.…
This paper develops a framework connecting discrete adjoint gradient-error analysis with an optimization method that uses directional error tolerances, and applies it to airfoil shape optimization governed by a conservative full-potential…
A variety of shooting methods for computing fully discrete time-periodic solutions of partial differential equations, including Newton-Krylov and optimization-based methods, are discussed and used to determine the periodic, compressible,…
We introduce a general-purpose method for optimising the mixing rate of advective fluid flows. An existing velocity field is perturbed in a $C^1$ neighborhood to maximize the mixing rate for flows generated by velocity fields in this…
Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they…