Related papers: Regularization lemmas and convergence in total var…
Distributionally robust optimization has emerged as an attractive way to train robust machine learning models, capturing data uncertainty and distribution shifts. Recent statistical analyses have proved that generalization guarantees of…
In this work, we investigate the convergence properties of the backward regularized Wasserstein proximal (BRWP) method for sampling a target distribution. The BRWP approach can be shown as a semi-implicit time discretization for a…
We introduce two algorithms for nonconvex regularized finite sum minimization, where typical Lipschitz differentiability assumptions are relaxed to the notion of relative smoothness. The first one is a Bregman extension of Finito/MISO,…
We investigate the properties of minimizers of one-dimensional variational problems when the Lagrangian has no higher smoothness than continuity. An elementary approximation result is proved, but it is shown that this cannot be in general…
Generalized linear models and the quasi-likelihood method extend the ordinary regression models to accommodate more general conditional distributions of the response. Nonparametric methods need no explicit parametric specification, and the…
Given a reference random variable, we study the solution of its Stein equation and obtain universal bounds on its first and second derivatives. We then extend the analysis of Nourdin and Peccati by bounding the Fortet-Mourier and…
Under general assumptions on the target distribution $p^\star$, we establish a sharp Lipschitz regularity theory for flow-matching vector fields and diffusion-model scores, with optimal dependence on time and dimension. As applications, we…
Although the \emph{residual method}, or \emph{constrained regularization}, is frequently used in applications, a detailed study of its properties is still missing. This sharply contrasts the progress of the theory of Tikhonov…
Regularization-based approaches for injecting constraints in Machine Learning (ML) were introduced to improve a predictive model via expert knowledge. We tackle the issue of finding the right balance between the loss (the accuracy of the…
We prove that Riemannian metrics with a uniform weak norm can be smoothed to having arbitrarily high regularity. This generalizes all previous smoothing results. As a consequence we obtain a generalization of Gromov's almost flat manifold…
Consider Ginibre's ensemble of $N \times N$ non-Hermitian random matrices in which all entries are independent complex Gaussians of mean zero and variance $\frac{1}{N}$. As $N \uparrow \infty$ the normalized counting measure of the…
We apply the recently introduced method of hermitization to study in the large $N$ limit non-hermitean random matrices that are drawn from a large class of circularly symmetric non-Gaussian probability distributions, thus extending the…
We propose a learning framework for graph kernels, which is theoretically grounded on regularizing optimal transport. This framework provides a novel optimal transport distance metric, namely Regularized Wasserstein (RW) discrepancy, which…
We study the problem of sampling from a distribution $\target$ using the Langevin Monte Carlo algorithm and provide rate of convergences for this algorithm in terms of Wasserstein distance of order $2$. Our result holds as long as the…
Wasserstein distributionally robust optimization (DRO) has recently achieved empirical success for various applications in operations research and machine learning, owing partly to its regularization effect. Although connection between…
In this paper, we propose a first second-order scheme based on arbitrary non-Euclidean norms, incorporated by Bregman distances. They are introduced directly in the Newton iterate with regularization parameter proportional to the square…
Fourier-Wiener transform of the formal expression for multiple self-intersection local time is described in terms of the integral, which is divergent on the diagonals. The method of regularization we use in this work related to…
This paper derives non-asymptotic error bounds for nonlinear stochastic approximation algorithms in the Wasserstein-$p$ distance. To obtain explicit finite-sample guarantees for the last iterate, we develop a coupling argument that compares…
We establish the first mathematically rigorous link between Bayesian, variational Bayesian, and ensemble methods. A key step towards this it to reformulate the non-convex optimisation problem typically encountered in deep learning as a…
We consider a semiparametric partly linear model identified by instrumental variables. We propose an estimation method that does not smooth on the instruments and we extend the Landweber-Fridman regularization scheme to the estimation of…