English
Related papers

Related papers: Indexability, concentration, and VC theory

200 papers

Concentration of distances in high dimension is an important factor for the development and design of stable and reliable data analysis algorithms. In this paper, we address the fundamental long-standing question about the concentration of…

Machine Learning · Computer Science 2026-04-01 Ivan Y. Tyukin , Bogdan Grechuk , Evgeny M. Mirkes , Alexander N. Gorban

High-dimensional data sets are commonly collected in many contemporary applications arising in various fields of scientific research. We present two views of finite samples in high dimensions: a probabilistic one and a nonprobabilistic one.…

Statistics Theory · Mathematics 2013-11-13 Jinchi Lv

Neuroscientists have recently shown that images that are difficult to find in visual search elicit similar patterns of firing across a population of recorded neurons. The $L^{1}$ distance between firing rate vectors associated with two…

Information Theory · Computer Science 2015-05-12 Nidhin Koshy Vaidhiyan , S. P. Arun , Rajesh Sundaresan

While there has been much interest in adapting conventional clustering procedures---and in higher dimensions, persistent homology methods---to directed networks, little is known about the convergence of such methods. In order to even…

Computational Geometry · Computer Science 2022-12-20 Samir Chowdhury , Facundo Mémoli

We derive concentration inequalities for functions of the empirical measure of large random matrices with infinitely divisible entries and, in particular, stable ones. We also give concentration results for some other functionals of these…

Probability · Mathematics 2007-06-13 Christian Houdré , Hua Xu

A first step when fitting multilevel models to continuous responses is to explore the degree of clustering in the data. Researchers fit variance-component models and then report the proportion of variation in the response that is due to…

Methodology · Statistics 2020-02-17 George Leckie , William Browne , Harvey Goldstein , Juan Merlo , Peter Austin

A concentration property of the functional ${-}\log f(X)$ is demonstrated, when a random vector X has a log-concave density f on $\mathbb{R}^n$. This concentration property implies in particular an extension of the Shannon-McMillan-Breiman…

Probability · Mathematics 2012-11-20 Sergey Bobkov , Mokshay Madiman

The Lipschitz constant of a neural network is connected to several important properties of the network such as its robustness and generalization. It is thus useful in many settings to estimate the Lipschitz constant of a model. Prior work…

Machine Learning · Computer Science 2026-03-02 Giannis Nikolentzos , Konstantinos Skianis

The article is devoted to the problem of inconsistency in the pairwise comparisons based prioritization methodology. The issue of "inconsistency" in this context has gained much attention in recent years. The literature provides us with a…

Artificial Intelligence · Computer Science 2015-10-22 Andrzej Z. Grzybowski

Clustering is an important phenomenon in turbulent flows laden with inertial particles. Although this process has been studied extensively, there are still open questions about both the fundamental physics and the reconciliation of…

Soft Condensed Matter · Physics 2025-10-09 Daniel Odens Mora , Alberto Aliseda , Alain Cartellier , Martin Obligado

We bound the number of nearly orthogonal vectors with fixed VC-dimension over $\setpm^n$. Our bounds are of interest in machine learning and empirical process theory and improve previous bounds by Haussler. The bounds are based on a simple…

Combinatorics · Mathematics 2011-02-18 Lee-Ad Gottlieb , Leonid , Kontorovich , Elchanan Mossel

The Natarajan dimension is a fundamental tool for characterizing multi-class PAC learnability, generalizing the Vapnik-Chervonenkis (VC) dimension from binary to multi-class classification problems. This work establishes upper bounds on…

Machine Learning · Statistics 2023-04-25 Ying Jin

Quantifying the similarity between two mathematical structures or datasets constitutes a particularly interesting and useful operation in several theoretical and applied problems. Aimed at this specific objective, the Jaccard index has been…

Machine Learning · Computer Science 2021-11-19 Luciano da F. Costa

Modern deep learning models have the ability to generate high-dimensional vectors whose similarity reflects semantic resemblance. Thus, similarity search, i.e., the operation of retrieving those vectors in a large collection that are…

Machine Learning · Computer Science 2024-04-04 Mariano Tepper , Ishwar Singh Bhati , Cecilia Aguerrebere , Mark Hildebrand , Ted Willke

This paper investigates score-based diffusion models when the underlying target distribution is concentrated on or near low-dimensional manifolds within the higher-dimensional space in which they formally reside, a common characteristic of…

Machine Learning · Computer Science 2025-01-03 Gen Li , Yuling Yan

Contrastive learning, a dominant self-supervised technique, emphasizes similarity in representations between augmentations of the same input and dissimilarity for different ones. Although low contrastive loss often correlates with high…

Machine Learning · Computer Science 2023-11-22 Yunzhe Zhang , Yao Lu , Qi Xuan

One of the goals of NASA funded project at IBM T. J. Watson Research Center was to build an index for similarity searching satellite images, which were characterized by high-dimensional feature image texture vectors. Reviewed is our effort…

Databases · Computer Science 2024-01-08 Alexander Thomasian

The Huge Object model of property testing [Goldreich and Ron, TheoretiCS 23] concerns properties of distributions supported on $\{0,1\}^n$, where $n$ is so large that even reading a single sampled string is unrealistic. Instead, query…

Data Structures and Algorithms · Computer Science 2024-12-04 Sourav Chakraborty , Eldar Fischer , Arijit Ghosh , Amit Levi , Gopinath Mishra , Sayantan Sen

We realized one dimensional disordered photonic structures by grouping high refractive index layers in clusters, randomly distributed in such structure. We have control on the maximum size of the cluster and on the ratio high-low refractive…

Materials Science · Physics 2016-09-21 Michele Bellingeri , Francesco Scotognella

Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between…

Machine Learning · Statistics 2017-09-29 Sebastijan Dumancic , Hendrik Blockeel