Related papers: Indexability, concentration, and VC theory
Clustering algorithms have significantly improved along with Deep Neural Networks which provide effective representation of data. Existing methods are built upon deep autoencoder and self-training process that leverages the distribution of…
Hyperuniformity refers to the suppression of density fluctuations at large scales. Typical for ordered systems, this property also emerges in several disordered physical and biological systems, where it is particularly relevant to…
Statistical learning theory provides bounds of the generalization gap, using in particular the Vapnik-Chervonenkis dimension and the Rademacher complexity. An alternative approach, mainly studied in the statistical physics literature, is…
In a complete metric space that is equipped with a doubling measure and supports a Poincar\'e inequality, we study strict subsets, i.e. sets whose variational capacity with respect to a larger reference set is finite, in the case $p=1$.…
A good measure of similarity between data points is crucial to many tasks in machine learning. Similarity and metric learning methods learn such measures automatically from data, but they do not scale well respect to the dimensionality of…
There has been growing interest in generalization performance of large multilayer neural networks that can be trained to achieve zero training error, while generalizing well on test data. This regime is known as 'second descent' and it…
We show that in a hierarchical clustering model the low-order statistics of the density and the peculiar velocity fields can all be modelled semianalytically for a given cosmology and an initial density perturbation power spectrum $P(k)$.…
Most dimensionality reduction methods employ frequency domain representations obtained from matrix diagonalization and may not be efficient for large datasets with relatively high intrinsic dimensions. To address this challenge, Correlated…
The paper concerns foundations of sensitivity and stability analysis in optimization and related areas, being primarily addressed truncated constrained systems. We consider general models, which are described by multifunctions between…
Collecting large-scale medical datasets with fine-grained annotations is time-consuming and requires experts. For this reason, weakly supervised learning aims at optimising machine learning models using weaker forms of annotations, such as…
We suggest that the curse of dimensionality affecting the similarity-based search in large datasets is a manifestation of the phenomenon of concentration of measure on high-dimensional structures. We prove that, under certain geometric…
We give a detailed asymptotic analysis of the profiles of random symmetric digital search trees, which are in close connection with the performance of the search complexity of random queries in such trees. While the expected profiles have…
Most Machine Learning (ML) methods, from clustering to classification, rely on a distance function to describe relationships between datapoints. For complex datasets it is hard to avoid making some arbitrary choices when defining a distance…
Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for…
In the present paper we obtain fully explicit large deviation inequalities for empirical processes indexed by a Vapnik--Chervonenkis class of sets (or functions). Furthermore we illustrate the importance of such results for the theory of…
We consider a random variable $X$ that takes values in a (possibly infinite-dimensional) topological vector space $\mathcal{X}$. We show that, with respect to an appropriate "normal distance" on $\mathcal{X}$, concentration inequalities for…
Clustering high-dimensional data is a critical challenge in machine learning due to the curse of dimensionality and the presence of noise. Traditional clustering algorithms often fail to capture the intrinsic structures in such data. This…
Vector representations and vector space modeling (VSM) play a central role in modern machine learning. We propose a novel approach to `vector similarity searching' over dense semantic representations of words and documents that can be…
This work continues the study of the relationship between sample compression schemes and statistical learning, which has been mostly investigated within the framework of binary classification. The central theme of this work is establishing…
Motivated by problems in high-dimensional statistics such as mixture modeling for classification and clustering, we consider the behavior of radial densities as the dimension increases. We establish a form of concentration of measure, and…