Related papers: Scale-Sensitive Shattering: Learnability and Evalu…
We study the problem of compression for the purpose of similarity identification, where similarity is measured by the mean square Euclidean distance between vectors. While the asymptotical fundamental limits of the problem - the minimal…
We study gapped scale-sensitive dimensions of a function class in both sequential and non-sequential settings. We demonstrate that covering numbers for any uniformly bounded class are controlled above by these gapped dimensions,…
We solve Talagrand's entropy problem: the L_2-covering numbers of every uniformly bounded class of functions are exponential in its shattering dimension. This extends Dudley's theorem on classes of {0,1}-valued functions, for which the…
Self-similar sets require a separation condition to admit a nice mathematical structure. The classical open set condition (OSC) is difficult to verify. Zerner proved that there is a positive and finite Hausdorff measure for a weaker…
Recently, learning algorithms motivated from sharpness of loss surface as an effective measure of generalization gap have shown state-of-the-art performances. Nevertheless, sharpness defined in a rigid region with a fixed radius, has a…
We study problem-dependent rates, i.e., generalization errors that scale near-optimally with the variance, the effective loss, or the gradient norms evaluated at the "best hypothesis." We introduce a principled framework dubbed "uniform…
How close are neural networks to the best they could possibly do? Standard benchmarks cannot answer this because they lack access to the true posterior p(y|x). We use class-conditional normalizing flows as oracles that make exact posteriors…
We prove that if an orientable 3-manifold $M$ admits a complete Riemannian metric whose scalar curvature is positive and has a subquadratic decay at infinity, then it decomposes as a (possibly infinite) connected sum of spherical manifolds…
For $n\in [-2,2]$ the $O(n)$ model on a random lattice has critical points to which a scaling behaviour characteristic of 2D gravity interacting with conformal matter fields with $c\in [-\infty,1]$ can be associated. Previously we have…
Calibration is a critical requirement for reliable probabilistic prediction, especially in high-risk applications. However, the theoretical understanding of which learning algorithms can simultaneously achieve high accuracy and good…
We analyze a family of supervised learning algorithms based on sample compression schemes that are stable, in the sense that removing points from the training set which were not selected for the compression set does not alter the resulting…
We study the adsorption problem of linear polymers, when the container of the polymer--solvent system is taken to be a member of the three dimensional Sierpinski gasket (SG) family of fractals. Members of the SG family are enumerated by an…
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity ($N$), datasize ($D$), and compute ($C$). However, existing theoretical explanations often rely on specific architectures or complex…
Algorithmic stability is a classical approach to understanding and analysis of the generalization error of learning algorithms. A notable weakness of most stability-based generalization bounds is that they hold only in expectation.…
We study a variant of Collaborative PAC Learning, in which we aim to learn an accurate classifier for each of the $n$ data distributions, while minimizing the number of samples drawn from them in total. Unlike in the usual collaborative…
We study the problem of efficient online multiclass linear classification with bandit feedback, where all examples belong to one of $K$ classes and lie in the $d$-dimensional Euclidean space. Previous works have left open the challenge of…
This paper studies quantum supervised learning for classical inference from quantum states. In this model, a learner has access to a set of labeled quantum samples as the training set. The objective is to find a quantum measurement that…
Learning from label proportions (LLP) is a generalization of supervised learning in which the training data is available as sets or bags of feature-vectors (instances) along with the average instance-label of each bag. The goal is to train…
We revisit the problem of characterising the complexity of Quantum PAC learning, as introduced by Bshouty and Jackson [SIAM J. Comput. 1998, 28, 1136-1153]. Several quantum advantages have been demonstrated in this setting, however, none…
We consider a class of nearest-neighbor weakly asymmetric mass conservative particle systems evolving on $\mathbb{Z}$, which includes zero-range and types of exclusion processes, starting from a perturbation of a stationary state. When the…