English
Related papers

Related papers: High-dimensional $p$-norms

200 papers

Asymptotic methods for hypothesis testing in high-dimensional data usually require the dimension of the observations to increase to infinity, often with an additional relationship between the dimension (say, $p$) and the sample size (say,…

Methodology · Statistics 2025-12-11 Ritabrata Karmakar , Joydeep Chowdhury , Subhajit Dutta , Marc G. Genton

We obtain concentration and large deviation for the sums of independent and identically distributed random variables with heavy-tailed distributions. Our concentration results are concerned with random variables whose distributions satisfy…

Probability · Mathematics 2022-07-27 Milad Bakhshizadeh , Arian Maleki , Victor H. de la Pena

Large deviation results are given for a class of perturbed nonhomogeneous Markov chains on finite state space which formally includes some stochastic optimization algorithms. Specifically, let {P_n} be a sequence of transition matrices on a…

Probability · Mathematics 2007-05-23 Zach Dietz , Sunder Sethuraman

We consider the problem of bounding large deviations for non-i.i.d. random variables that are allowed to have arbitrary dependencies. Previous works typically assumed a specific dependence structure, namely the existence of independent…

Probability · Mathematics 2018-11-06 Christoph H. Lampert , Liva Ralaivola , Alexander Zimin

We revisit the problem of condensation for independent, identically distributed random variables with a power-law tail, conditioned by the value of their sum. For large values of the sum, and for a large number of summands, a condensation…

Statistical Mechanics · Physics 2022-03-03 Claude Godrèche

High-dimensional multivariate time series are challenging due to the dependent and high-dimensional nature of the data, but in many applications there is additional structure that can be exploited to reduce computing time along with…

Methodology · Statistics 2020-03-13 Michael Schweinberger , Sergii Babkin , Katherine Ensor

The notion of interpolation and extrapolation is fundamental in various fields from deep learning to function approximation. Interpolation occurs for a sample $x$ whenever this sample falls inside or on the boundary of the given dataset's…

Machine Learning · Computer Science 2021-11-02 Randall Balestriero , Jerome Pesenti , Yann LeCun

We suggest that the curse of dimensionality affecting the similarity-based search in large datasets is a manifestation of the phenomenon of concentration of measure on high-dimensional structures. We prove that, under certain geometric…

Information Retrieval · Computer Science 2009-11-17 Vladimir Pestov

Let L be a positive line bundle over a projective complex manifold X. Consider the space of holomorphic sections of the tensor power of order p of L. The determinant of a basis of this space, together with some given probability measure on…

Complex Variables · Mathematics 2016-03-14 Tien-Cuong Dinh , Viet-Anh Nguyen

The advent of modern technology, permitting the measurement of thousands of characteristics simultaneously, has given rise to floods of data characterized by many large or even huge datasets. This new paradigm presents extraordinary…

Methodology · Statistics 2019-02-14 A. M. Pires , J. A. Branco

In repeated Measure Designs with multiple groups, the primary purpose is to compare different groups in various aspects. For several reasons, the number of measurements and therefore the dimension of the observation vectors can depend on…

Statistics Theory · Mathematics 2022-07-20 Paavo Sattler , Markus Pauly

In this article we prove three fundamental types of limit theorems for the $q$-norm of random vectors chosen at random in an $\ell_p^n$-ball in high dimensions. We obtain a central limit theorem, a moderate deviations as well as a large…

Probability · Mathematics 2019-06-11 Zakhar Kabluchko , Joscha Prochno , Christoph Thaele

In this paper, we show the central limit theorem for the logarithmic determinant of the sample correlation matrix $\mathbf{R}$ constructed from the $(p\times n)$-dimensional data matrix $\mathbf{X}$ containing independent and identically…

Probability · Mathematics 2023-02-27 Johannes Heiny , Nestor Parolya

We derive sharp upper and lower bounds for the pointwise concentration function of the maximum statistic of $d$ identically distributed real-valued random variables. Our first main result places no restrictions either on the common marginal…

Statistics Theory · Mathematics 2025-08-04 Matias D. Cattaneo , Ricardo P. Masini , William G. Underwood

We study the problem of mean estimation for high-dimensional distributions, assuming access to a statistical query oracle for the distribution. For a normed space $X = (\mathbb{R}^d, \|\cdot\|_X)$ and a distribution supported on vectors $x…

Data Structures and Algorithms · Computer Science 2019-02-08 Jerry Li , Aleksandar Nikolov , Ilya Razenshteyn , Erik Waingarten

Different types of two- and three-dimensional representations of a finite metric space are studied that focus on the accurate representation of the linear order among the distances rather than their actual values. Lower and upper bounds for…

Combinatorics · Mathematics 2007-05-23 Jobst Heitzig

Most Machine Learning (ML) methods, from clustering to classification, rely on a distance function to describe relationships between datapoints. For complex datasets it is hard to avoid making some arbitrary choices when defining a distance…

Machine Learning · Statistics 2016-07-04 Gina Gruenhage , Manfred Opper , Simon Barthelme

Given $p \in (0,1)$, we let $Q_p= Q_p^d$ be the random subgraph of the $d$-dimensional hypercube $Q^d$ where edges are present independently with probability $p$. It is well known that, as $d \rightarrow \infty$, if $p>\frac12$ then with…

Combinatorics · Mathematics 2021-01-05 Colin McDiarmid , Alex Scott , Paul Withers

Consider that the coordinates of $N$ points are randomly generated along the edges of a $d$-dimensional hypercube (random point problem). The probability that an arbitrary point is the $m$th nearest neighbor to its own $n$th nearest…

Disordered Systems and Neural Networks · Physics 2007-05-23 Cesar Augusto Sangaletti Tercariol , Felipe de Mouta Kiipper , Alexandre Souto Martinez

In this work, we study distance metric learning (DML) for high dimensional data. A typical approach for DML with high dimensional data is to perform the dimensionality reduction first before learning the distance metric. The main…

Machine Learning · Computer Science 2015-09-16 Qi Qian , Rong Jin , Lijun Zhang , Shenghuo Zhu