Related papers: High-dimensional $p$-norms
We provide some asymptotic theory for the largest eigenvalues of a sample covariance matrix of a p-dimensional time series where the dimension p = p_n converges to infinity when the sample size n increases. We give a short overview of the…
Thanks to its favorable properties, the multivariate normal distribution is still largely employed for modeling phenomena in various scientific fields. However, when the number of components $p$ is of the same asymptotic order as the sample…
Let ${X}_{k}=(x_{k1}, \cdots, x_{kp})', k=1,\cdots,n$, be a random sample of size $n$ coming from a $p$-dimensional population. For a fixed integer $m\geq 2$, consider a hypercubic random tensor $\mathbf{{T}}$ of $m$-th order and rank $n$…
Central limit theorems are established for the sum, over a spatial region, of observations from a linear process on a $d$-dimensional lattice. This region need not be rectangular, but can be irregularly-shaped. Separate results are…
We perform a deeper analysis of an axiomatic approach to the concept of intrinsic dimension of a dataset proposed by us in the IJCNN'07 paper (arXiv:cs/0703125). The main features of our approach are that a high intrinsic dimension of a…
Consider multiple sums $S_n$ on the $d$-dimensional integer grid,which are generated by i.i.d.\ random variables with a positive expectation. We prove the strong law of large numbers, the law of the iterated logarithm and the distributional…
We introduce $p$-uniformity to characterize the scaling of density fluctuations in spatial random systems in $\mathbb{R}^d$, ranging from hyperfluctuation to stealthy hyperuniformity. Our central theorem establishes sufficient conditions to…
Many high dimensional integrals can be reduced to the problem of finding the relative measures of two sets. Often one set will be exponentially larger than the other, making it difficult to compare the sizes. A standard method of dealing…
We derive exact statistical properties of a class of recursive fragmentation processes. We show that introducing a fragmentation probability 0<p<1 leads to a purely algebraic size distribution in one dimension, P(x) ~ x^{-2p}. In d…
In this article, we propose some two-sample tests based on ball divergence and investigate their high dimensional behavior. First, we study their behavior for High Dimension, Low Sample Size (HDLSS) data, and under appropriate regularity…
We establish central and local limit theorems for the number of vertices in the largest component of a random $d$-uniform hypergraph $\hnp$ with edge probability $p=c/\binnd$, where $(d-1)^{-1}+\eps<c<\infty$. The proof relies on a new,…
The phenomenon of superconvergence is proved for all freely infinitely divisible distributions. Precisely, suppose that the partial sums of a sequence of free identically distributed, infinitesimal random variables converge in distribution…
We consider the problem of finding high dimensional approximate nearest neighbors. Suppose there are d independent rare features, each having its own independent statistics. A point x will have x_{i}=0 denote the absence of feature i, and…
Central limit theorems (CLTs) for high-dimensional random vectors with dimension possibly growing with the sample size have received a lot of attention in the recent times. Chernozhukov et al. (2017) proved a Berry--Esseen type result for…
The study of high-dimensional distributions is of interest in probability theory, statistics and asymptotic convex geometry, where the object of interest is the uniform distribution on a convex set in high dimensions. The $\ell^p$ spaces…
This article proposes a new approach to modeling high-dimensional time series by treating a $p$-dimensional time series as a nonsingular linear transformation of certain common factors and idiosyncratic components. Unlike the approximate…
Let $p$ be an unknown and arbitrary probability distribution over $[0,1)$. We consider the problem of {\em density estimation}, in which a learning algorithm is given i.i.d. draws from $p$ and must (with high probability) output a…
Let $X$ be an isotropic random vector in $R^d$ that satisfies that for every $v \in S^{d-1}$, $\|<X,v>\|_{L_q} \leq L \|<X,v>\|_{L_p}$ for some $q \geq 2p$. We show that for $0<\varepsilon<1$, a set of $N = c(p,q,\varepsilon) d$ random…
This paper presents a unified framework, integrating information theory and statistical mechanics, to connect metric failure in high-dimensional data with emergence in complex systems. We propose the "Information Dilution Theorem,"…
These are lecture notes based on the first part of a course on 'Mathematical Data Science', which I taught to final year BSc students in the UK in 2019-2020. Topics include: concentration of measure in high dimensions; Gaussian random…