English
Related papers

Related papers: Improving Generalization Bounds for VC Classes Usi…

200 papers

We establish upper and lower bounds with matching leading terms for tails of weighted sums of two-sided exponential random variables. This extends Janson's recent results for one-sided exponentials.

Probability · Mathematics 2025-01-28 Jiawei Li , Tomasz Tkocz

Despite its success in a wide range of applications, characterizing the generalization properties of stochastic gradient descent (SGD) in non-convex deep learning problems is still an important challenge. While modeling the trajectories of…

Machine Learning · Statistics 2022-01-12 Umut Şimşekli , Ozan Sener , George Deligiannidis , Murat A. Erdogdu

In the real world, the frequency of occurrence of objects is naturally skewed forming long-tail class distributions, which results in poor performance on the statistically rare classes. A promising solution is to mine tail-class examples to…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Gursimran Singh , Lingyang Chu , Lanjun Wang , Jian Pei , Qi Tian , Yong Zhang

This work prepares new probability bounds for sums of random, independent, Hermitian tensors. These probability bounds characterize large-deviation behavior of the extreme eigenvalue of the sums of random tensors. We extend Lapalace…

Probability · Mathematics 2021-01-01 Shih Yu Chang

This dissertation studies a fundamental open challenge in deep learning theory: why do deep networks generalize well even while being overparameterized, unregularized and fitting the training data to zero error? In the first part of the…

Machine Learning · Computer Science 2021-10-19 Vaishnavh Nagarajan

Let $X$ be a centered random vector in a finite dimensional real inner product space $\mathcal{E}$. For a subset $C$ of the ambient vector space $V$ of $\mathcal{E}$ and $x,\,y\in V$, write $x\preceq_C y$ if $y-x\in C$. When $C$ is a closed…

Probability · Mathematics 2023-07-10 Nicola Apollonio

We give a new proof of VC bounds where we avoid the use of symmetrization and use a shadow sample of arbitrary size. We also improve on the variance term. This results in better constants, as shown on numerical examples. Moreover our bounds…

Statistics Theory · Mathematics 2007-06-13 Olivier Catoni

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most…

Machine Learning · Computer Science 2026-01-28 Hongxu Chen , Ke Wei , Xiaoming Yuan , Luo Luo

The best practical techniques for exact solution of instances of the constrained maximum-entropy sampling problem, a discrete-optimization problem arising in the design of experiments, are via a branch-and-bound framework, working with a…

Optimization and Control · Mathematics 2024-02-19 Zhongzhu Chen , Marcia Fampa , Jon Lee

This paper investigates the efficiency of the K-fold cross-validation (CV) procedure and a debiased version thereof as a means of estimating the generalization risk of a learning algorithm. We work under the general assumption of uniform…

Statistics Theory · Mathematics 2023-06-13 Anass Aghbalou , François Portier , Anne Sabourin

We present new algorithms for online convex optimization over unbounded domains that obtain parameter-free regret in high-probability given access only to potentially heavy-tailed subgradient estimates. Previous work in unbounded domains…

Machine Learning · Statistics 2023-02-28 Jiujia Zhang , Ashok Cutkosky

We analyze the sample complexity of full-batch Gradient Descent (GD) in the setup of non-smooth Stochastic Convex Optimization. We show that the generalization error of GD, with common choice of hyper-parameters, can be $\tilde \Theta(d/m +…

Machine Learning · Computer Science 2024-04-12 Roi Livni

We study large deviation upper bounds and mean-squared error (MSE) guarantees of a general framework of nonlinear stochastic gradient methods in the online setting, in the presence of heavy-tailed noise. Unlike existing works that rely on…

Machine Learning · Computer Science 2025-03-25 Aleksandar Armacki , Shuhua Yu , Dragana Bajovic , Dusan Jakovetic , Soummya Kar

Background: Deep learning models are typically trained using stochastic gradient descent or one of its variants. These methods update the weights using their gradient, estimated from a small fraction of the training data. It has been…

Machine Learning · Statistics 2018-01-03 Elad Hoffer , Itay Hubara , Daniel Soudry

Using tail bounds, we introduce a new probabilistic condition for function estimation in stochastic derivative-free optimization which leads to a reduction in the number of samples and eases algorithmic analyses. Moreover, we develop simple…

Optimization and Control · Mathematics 2023-06-16 Francesco Rinaldi , Luis Nunes Vicente , Damiano Zeffiro

In this paper we establish a new margin-based generalization bound for voting classifiers, refining existing results and yielding tighter generalization guarantees for widely used boosting algorithms such as AdaBoost (Freund and Schapire,…

Machine Learning · Computer Science 2025-06-04 Mikael Møller Høgsgaard , Kasper Green Larsen

We study the upper tail of the number of arithmetic progressions of a given length in a random subset of {1,...,n}, establishing exponential bounds which are best possible up to constant factors in the exponent. The proof also extends to…

Combinatorics · Mathematics 2017-12-12 Lutz Warnke

Gradient clipping is a commonly used technique to stabilize the training process of neural networks. A growing body of studies has shown that gradient clipping is a promising technique for dealing with the heavy-tailed behavior that emerged…

Machine Learning · Computer Science 2023-07-26 Shaojie Li , Yong Liu

We consider non-convex stochastic optimization using first-order algorithms for which the gradient estimates may have heavy tails. We show that a combination of gradient clipping, momentum, and normalized gradient descent yields convergence…

Machine Learning · Computer Science 2021-11-10 Ashok Cutkosky , Harsh Mehta

Variational inference (VI) is widely used for approximate inference in Bayesian machine learning. In addition to this practical success, generalization bounds for variational inference and related algorithms have been developed, mostly…

Machine Learning · Computer Science 2025-02-19 Yadi Wei , Roni Khardon
‹ Prev 1 4 5 6 7 8 10 Next ›