Related papers: Uniform Convergence with Square-Root Lipschitz Los…
We prove a universality theorem for learning with random features. Our result shows that, in terms of training and generalization errors, a random feature model with a nonlinear activation function is asymptotically equivalent to a…
The strong convergence of Euler approximations of stochastic delay differential equations is proved under general conditions. The assumptions on drift and diffusion coefficients have been relaxed to include polynomial growth and only…
For an inverse coefficient problem of determining a state-varying factor in the corresponding Hamiltonian for a mean field game system, we prove the global Lipschitz stability by spatial data of one component and interior data in an…
We consider the fundamental problem of ReLU regression, where the goal is to output the best fitting ReLU with respect to square loss given access to draws from some unknown distribution. We give the first efficient, constant-factor…
Gaussian universality results assert that the properties of many estimators remain unchanged when the input data are replaced by Gaussians. Such results have gained popularity in high-dimensional statistics and machine learning, as…
Theoretical estimates of the convergence rate of many well-known gradient-type optimization methods are based on quadratic interpolation, provided that the Lipschitz condition for the gradient is satisfied. In this article we obtain a…
We establish an excess risk bound of O(H R_n^2 + R_n \sqrt{H L*}) for empirical risk minimization with an H-smooth loss function and a hypothesis class with Rademacher complexity R_n, where L* is the best risk achievable by the hypothesis…
The purpose of this note is to provide an optimal rate of convergence in the vanishing viscosity regime for first-order Hamilton-Jacobi equations with uniformly convex Hamiltonian. We prove that for a globally Lipschitz-continuous and…
Data valuation quantifies data importance, but existing methods cannot ensure validity in a single training process. The neural dynamic data valuation (NDDV) method [3] addresses this limitation. Based on NDDV, we are the first to explore…
The self-concordant-like property of a smooth convex function is a new analytical structure that generalizes the self-concordant notion. While a wide variety of important applications feature the self-concordant-like property, this concept…
In a previous paper it was shown that the Forward Euler method applied to differential inclusions where the right-hand side is a Lipschitz continuous set-valued function with uniformly bounded, compact values, converges with rate one. The…
We describe algorithms for finding the regression of t, a sequence of values, to the closest sequence s by mean squared error, so that s is always increasing (isotonicity) and so the values of two consecutive points do not increase by too…
Many machine learning techniques sacrifice convenient computational structures to gain estimation robustness and modeling flexibility. However, by exploring the modeling structures, we find these "sacrifices" do not always require more…
Symmetry analysis can provide a suitable change of variables, i.e., in geometric terms, a suitable diffeomorphism that simplifies the given direction field, which can help significantly in solving or studying differential equations. Roughly…
We study the generalization performance of unregularized gradient methods for separable linear classification. While previous work mostly deal with the binary case, we focus on the multiclass setting with $k$ classes and establish novel…
In this paper we introduce a randomized version of the backward Euler method, that is applicable to stiff ordinary differential equations and nonlinear evolution equations with time-irregular coefficients. In the finite-dimensional case, we…
We consider solutions satisfying the zero Neumann boundary condition and a linearized mean field game equation in $\Omega \times (0,T)$ whose principal coefficients depend on the time and spatial variables with general Hamiltonian, where…
This manuscript bridges nonparametric smoothness-based and shape-restricted estimation, which may appear as two disjoint paradigms in the field. The proposed approach is motivated by a conceptually simple observation: every Lipschitz…
We develop and analyze stochastic optimization algorithms for problems in which the expected loss is strongly convex, and the optimum is (approximately) sparse. Previous approaches are able to exploit only one of these two structures,…
This paper is a part of our programme to generalise the Hardy-Littlewood method to handle systems of linear questions in primes. This programme is laid out in our paper Linear Equations in Primes [LEP], which accompanies this submission. In…