English
Related papers

Related papers: Uniform Convergence with Square-Root Lipschitz Los…

200 papers

We prove a universality theorem for learning with random features. Our result shows that, in terms of training and generalization errors, a random feature model with a nonlinear activation function is asymptotically equivalent to a…

Information Theory · Computer Science 2022-11-01 Hong Hu , Yue M. Lu

The strong convergence of Euler approximations of stochastic delay differential equations is proved under general conditions. The assumptions on drift and diffusion coefficients have been relaxed to include polynomial growth and only…

Probability · Mathematics 2013-03-07 Chaman Kumar , Sotirios Sabanis

For an inverse coefficient problem of determining a state-varying factor in the corresponding Hamiltonian for a mean field game system, we prove the global Lipschitz stability by spatial data of one component and interior data in an…

Analysis of PDEs · Mathematics 2023-07-11 Oleg Imanuvilov , Masahiro Yamamoto

We consider the fundamental problem of ReLU regression, where the goal is to output the best fitting ReLU with respect to square loss given access to draws from some unknown distribution. We give the first efficient, constant-factor…

Machine Learning · Computer Science 2020-09-30 Ilias Diakonikolas , Surbhi Goel , Sushrut Karmalkar , Adam R. Klivans , Mahdi Soltanolkotabi

Gaussian universality results assert that the properties of many estimators remain unchanged when the input data are replaced by Gaussians. Such results have gained popularity in high-dimensional statistics and machine learning, as…

Probability · Mathematics 2025-12-03 Kevin Han Huang , Morgane Austern , Peter Orbanz

Theoretical estimates of the convergence rate of many well-known gradient-type optimization methods are based on quadratic interpolation, provided that the Lipschitz condition for the gradient is satisfied. In this article we obtain a…

Optimization and Control · Mathematics 2018-12-18 Fedor S. Stonyakin

We establish an excess risk bound of O(H R_n^2 + R_n \sqrt{H L*}) for empirical risk minimization with an H-smooth loss function and a hypothesis class with Rademacher complexity R_n, where L* is the best risk achievable by the hypothesis…

Machine Learning · Computer Science 2012-11-27 Nathan Srebro , Karthik Sridharan , Ambuj Tewari

The purpose of this note is to provide an optimal rate of convergence in the vanishing viscosity regime for first-order Hamilton-Jacobi equations with uniformly convex Hamiltonian. We prove that for a globally Lipschitz-continuous and…

Analysis of PDEs · Mathematics 2025-06-17 Louis-Pierre Chaintron , Samuel Daudin

Data valuation quantifies data importance, but existing methods cannot ensure validity in a single training process. The neural dynamic data valuation (NDDV) method [3] addresses this limitation. Based on NDDV, we are the first to explore…

Machine Learning · Computer Science 2025-12-19 Zhangyong Liang , Huanhuan Gao , Ji Zhang

The self-concordant-like property of a smooth convex function is a new analytical structure that generalizes the self-concordant notion. While a wide variety of important applications feature the self-concordant-like property, this concept…

Optimization and Control · Mathematics 2018-01-23 Quoc Tran-Dinh , Yen-Huan Li , Volkan Cevher

In a previous paper it was shown that the Forward Euler method applied to differential inclusions where the right-hand side is a Lipschitz continuous set-valued function with uniformly bounded, compact values, converges with rate one. The…

Numerical Analysis · Mathematics 2009-02-02 Mattias Sandberg

We describe algorithms for finding the regression of t, a sequence of values, to the closest sequence s by mean squared error, so that s is always increasing (isotonicity) and so the values of two consecutive points do not increase by too…

Data Structures and Algorithms · Computer Science 2009-12-31 Pankaj K. Agarwal , Jeff M. Phillips , Bardia Sadri

Many machine learning techniques sacrifice convenient computational structures to gain estimation robustness and modeling flexibility. However, by exploring the modeling structures, we find these "sacrifices" do not always require more…

Machine Learning · Computer Science 2019-04-16 Xingguo Li , Haoming Jiang , Jarvis Haupt , Raman Arora , Han Liu , Mingyi Hong , Tuo Zhao

Symmetry analysis can provide a suitable change of variables, i.e., in geometric terms, a suitable diffeomorphism that simplifies the given direction field, which can help significantly in solving or studying differential equations. Roughly…

Classical Analysis and ODEs · Mathematics 2020-10-02 Eszter Gselmann , Gábor Horváth

We study the generalization performance of unregularized gradient methods for separable linear classification. While previous work mostly deal with the binary case, we focus on the multiclass setting with $k$ classes and establish novel…

Machine Learning · Computer Science 2025-05-29 Matan Schliserman , Tomer Koren

In this paper we introduce a randomized version of the backward Euler method, that is applicable to stiff ordinary differential equations and nonlinear evolution equations with time-irregular coefficients. In the finite-dimensional case, we…

Numerical Analysis · Mathematics 2022-05-10 Monika Eisenmann , Mihály Kovács , Raphael Kruse , Stig Larsson

We consider solutions satisfying the zero Neumann boundary condition and a linearized mean field game equation in $\Omega \times (0,T)$ whose principal coefficients depend on the time and spatial variables with general Hamiltonian, where…

Analysis of PDEs · Mathematics 2023-04-14 Oleg Imanuvilov , Hongyu Liu , Masahiro Yamamoto

This manuscript bridges nonparametric smoothness-based and shape-restricted estimation, which may appear as two disjoint paradigms in the field. The proposed approach is motivated by a conceptually simple observation: every Lipschitz…

Methodology · Statistics 2026-05-22 Kenta Takatsu , Tianyu Zhang , Arun Kumar Kuchibhotla

We develop and analyze stochastic optimization algorithms for problems in which the expected loss is strongly convex, and the optimum is (approximately) sparse. Previous approaches are able to exploit only one of these two structures,…

Machine Learning · Statistics 2012-07-19 Alekh Agarwal , Sahand Negahban , Martin J. Wainwright

This paper is a part of our programme to generalise the Hardy-Littlewood method to handle systems of linear questions in primes. This programme is laid out in our paper Linear Equations in Primes [LEP], which accompanies this submission. In…

Number Theory · Mathematics 2011-11-09 Ben Green , Terence Tao
‹ Prev 1 8 9 10 Next ›