English
Related papers

Related papers: Spectrally-normalized margin bounds for neural net…

200 papers

Self-normalized martingale inequalities lie at the heart of confidence ellipsoids for online least squares and, more broadly, many bandit and reinforcement-learning results. Yet existing vector and scalar results typically rely on bounded…

Machine Learning · Statistics 2026-05-05 Fan Chen , Jian Qian , Alexander Rakhlin , Nikita Zhivotovskiy

We present a brief survey of existing mistake bounds and introduce novel bounds for the Perceptron or the kernel Perceptron algorithm. Our novel bounds generalize beyond standard margin-loss type bounds, allow for any convex and Lipschitz…

Machine Learning · Computer Science 2013-07-24 Mehryar Mohri , Afshin Rostamizadeh

Stochastic gradient descent (SGD) and adaptive gradient methods, such as Adam and RMSProp, have been widely used in training deep neural networks. We empirically show that while the difference between the standard generalization performance…

Machine Learning · Computer Science 2023-11-30 Avery Ma , Yangchen Pan , Amir-massoud Farahmand

This paper follows up on a recent work of Neu et al. (2021) and presents some new information-theoretic upper bounds for the generalization error of machine learning models, such as neural networks, trained with SGD. We apply these bounds…

Machine Learning · Computer Science 2022-03-22 Ziqiao Wang , Yongyi Mao

Despite recent success, state-of-the-art learning-based models remain highly vulnerable to input changes such as adversarial examples. In order to obtain certifiable robustness against such perturbations, recent work considers…

Machine Learning · Computer Science 2023-09-13 Max Losch , David Stutz , Bernt Schiele , Mario Fritz

The Lipschitz constant is a key measure for certifying the robustness of neural networks to input perturbations. However, computing the exact constant is NP-hard, and standard approaches to estimate the Lipschitz constant involve solving a…

Machine Learning · Computer Science 2026-04-14 Yuezhu Xu , S. Sivaranjani

We initiate the study of nonsmooth optimization problems under bounded local subgradient variation, which postulates bounded difference between (sub)gradients in small local regions around points, in either average or maximum sense. The…

Optimization and Control · Mathematics 2024-11-05 Jelena Diakonikolas , Cristóbal Guzmán

We show generalisation error bounds for deep learning with two main improvements over the state of the art. (1) Our bounds have no explicit dependence on the number of classes except for logarithmic factors. This holds even when formulating…

Machine Learning · Computer Science 2021-02-23 Antoine Ledent , Waleed Mustafa , Yunwen Lei , Marius Kloft

We study non-convex empirical risk minimization for learning halfspaces and neural networks. For loss functions that are $L$-Lipschitz continuous, we present algorithms to learn halfspaces and multi-layer neural networks that achieve…

Machine Learning · Computer Science 2015-11-26 Yuchen Zhang , Jason D. Lee , Martin J. Wainwright , Michael I. Jordan

Let $n\ge2$ and $\Omega$ be a bounded Lipschitz domain in $\mathbb{R}^n$. In this article, the authors investigate global (weighted) estimates for the gradient of solutions to Robin boundary value problems of second order elliptic equations…

Analysis of PDEs · Mathematics 2020-03-18 Sibei Yang , Dachun Yang , Wen Yuan

Graph neural networks (GNNs) have recently been demonstrated to perform well on a variety of network-based tasks such as decentralized control and resource allocation, and provide computationally efficient methods for these tasks which have…

Machine Learning · Computer Science 2021-12-15 Raghu Arghal , Eric Lei , Shirin Saeedi Bidokhti

We deal with homogeneous Dirichlet and Neumann boundary-value problems for anisotropic elliptic operators of p-Laplace type. They emerge as Euler-Lagrange equations of integral functionals of the Calculus of Variations built upon possibly…

Analysis of PDEs · Mathematics 2025-10-28 Carlo Alberto Antonini , Andrea Cianchi

It has been experimentally observed in recent years that multi-layer artificial neural networks have a surprising ability to generalize, even when trained with far more parameters than observations. Is there a theoretical basis for this?…

Machine Learning · Statistics 2018-09-19 Andrew R. Barron , Jason M. Klusowski

This paper demonstrates the robustness of Lipschitz-regularized $\alpha$-divergences as objective functionals in generative modeling, showing they enable stable learning across a wide range of target distributions with minimal assumptions.…

Machine Learning · Statistics 2025-09-09 Ziyu Chen , Hyemin Gu , Markos A. Katsoulakis , Luc Rey-Bellet , Wei Zhu

Decentralized optimization has become a fundamental tool for large-scale learning systems; however, most existing methods rely on the classical Lipschitz smoothness assumption, which is often violated in problems with rapidly varying…

Optimization and Control · Mathematics 2026-01-08 Yanan Bo , Yongqiang Wang

In this paper we consider Deep Neural Networks (DNNs) with a smooth activation function as surrogates for high-dimensional functions that are somewhat smooth but costly to evaluate. We consider the standard (non-periodic) DNNs as well as…

Numerical Analysis · Mathematics 2026-03-04 Alexander Keller , Frances Y. Kuo , Dirk Nuyens , Ian H. Sloan

In this paper, we study the implicit regularization of the gradient descent algorithm in homogeneous neural networks, including fully-connected and convolutional neural networks with ReLU or LeakyReLU activations. In particular, we study…

Machine Learning · Computer Science 2021-01-01 Kaifeng Lyu , Jian Li

We investigate how the final parameters found by stochastic gradient descent are influenced by over-parameterization. We generate families of models by increasing the number of channels in a base network, and then perform a large…

Machine Learning · Computer Science 2019-05-10 Daniel S. Park , Jascha Sohl-Dickstein , Quoc V. Le , Samuel L. Smith

Trace norm regularization is a popular method of multitask learning. We give excess risk bounds with explicit dependence on the number of tasks, the number of examples per task and properties of the data distribution. The bounds are…

Machine Learning · Statistics 2013-01-15 Andreas Maurer , Massimiliano Pontil

Neural ordinary differential equations (neural ODEs) are a popular type of deep learning model that operate with continuous-depth architectures. To assess how well such models perform on unseen data, it is crucial to understand their…

Machine Learning · Computer Science 2025-08-27 Madhusudan Verma , Manoj Kumar