Related papers: Kurdyka-{\L}ojasiewicz exponent via Hadamard param…
We study the optimization of non-convex functions that are not necessarily smooth (gradient and/or Hessian are Lipschitz) using first order methods. Smoothness is a restrictive assumption in machine learning in both theory and practice,…
We give first-order asymptotic expansions for the resolvent and Hadamard-type formulas for the eigenvalue curves of one-parameter families of canonically symplectic operators. We allow for parameter dependence in the boundary conditions,…
The statistical behaviour of a product of independent, identically distributed random matrices in $\text{SL}(2,{\mathbb R})$ is encoded in the generalised Lyapunov exponent $\Lambda$; this is a function whose value at the complex number $2…
Aim of this paper is to prove the second order differentiation formula for $H^{2,2}$ functions along geodesics in $RCD^*(K,N)$ spaces with $N < \infty$. This formula is new even in the context of Alexandrov spaces, where second order…
Overdetermined systems of first kind integral equations appear in many applications. When the right-hand side is discretized, the resulting finite-data problem is ill-posed and admits infinitely many solutions. We propose a numerical method…
This paper is concerned with a modified entropy method to establish the large-time convergence towards the (unique) steady state, for kinetic Fokker-Planck equations with non-quadratic confinement potentials in whole space. We extend…
It is well-known that saturated output observations are prevalent in various practical systems and that the $\ell_1$-norm is more robust than the $\ell_2$-norm-based parameter estimation. Unfortunately, adaptive identification based on both…
We solve the Anderson localization problem on a two-leg ladder by the Fokker-Planck equation approach. The solution is exact in the weak disorder limit at a fixed inter-chain coupling. The study is motivated by progress in investigating the…
We provide a simple and flexible framework for designing differentially private algorithms to find approximate stationary points of non-convex loss functions. Our framework is based on using a private approximate risk minimizer to "warm…
Large over-parametrized models learned via stochastic gradient descent (SGD) methods have become a key element in modern machine learning. Although SGD methods are very effective in practice, most theoretical analyses of SGD suggest slower…
At the forefront of state-of-the-art human alignment methods are preference optimization methods (*PO). Prior research has often concentrated on identifying the best-performing method, typically involving a grid search over hyperparameters,…
In this paper, we study a second-order accurate and linear numerical scheme for the nonlocal Cahn-Hilliard equation. The scheme is established by combining a modified Crank-Nicolson approximation and the Adams-Bashforth extrapolation for…
The success of deep learning is due, to a large extent, to the remarkable effectiveness of gradient-based optimization methods applied to large neural networks. The purpose of this work is to propose a modern view and a general mathematical…
Let $L$ be a one-to-one operator of type $\omega$ in $L^2(\mathbb{R}^n)$, with $\omega\in[0,\,\pi/2)$, which has a bounded holomorphic functional calculus and satisfies the Davies-Gaffney estimates. Let $p(\cdot):\ \mathbb{R}^n\to(0,\,1]$…
We consider the nonconvex minimization problem, with quartic objective function, that arises in the exact recovery of a configuration matrix $P\in \R^{nd}$ of $n$ points when a Euclidean distance matrix, \EDMp, is given with embedding…
Stochastic differentiable approximation schemes are widely used for solving high dimensional problems. Most of existing methods satisfy some desirable properties, including conditional descent inequalities, and almost sure (a.s.)…
Existing example-based prediction explanation methods often bridge test and training data points through the model's parameters or latent representations. While these methods offer clues to the causes of model predictions, they often…
In a real Hilbert space setting, we reconsider the classical Arrow-Hurwicz differential system in view of solving linearly constrained convex minimization problems. We investigate the asymptotic properties of the differential system and…
Dynamic Mode Decomposition (DMD) and its variants, such as extended DMD (EDMD), are broadly used to fit simple linear models to dynamical systems known from observable data. As DMD methods work well in several situations but perform poorly…
We consider approximations to the solutions of differential Riccati equations in the context of linear quadratic regulator problems, where the state equation is governed by a multiscale operator. Similarly to elliptic and parabolic…