Related papers: Entropy-based Bounds on Dimension Reduction in L_1
Why does the low dimensionality of representations, typically $d\approx 1000$, not prevent modern embedding-based retrieval models from scaling to billions, or even trillions, of data points? To answer this question, we study maximal-margin…
In this note we discuss a common misconception, namely that embeddings are always used to reduce the dimensionality of the item space. We show that when we measure dimensionality in terms of information entropy then the embedding of sparse…
For every e>0, any subset of R^n with Hausdorff dimension larger than (1-e)n must have ultrametric distortion larger than 1/(4e).
We solve Talagrand's entropy problem: the L_2-covering numbers of every uniformly bounded class of functions are exponential in its shattering dimension. This extends Dudley's theorem on classes of {0,1}-valued functions, for which the…
Let (lambda_d)(p) be the p monomer-dimer entropy on the d-dimensional integer lattice Z^d, where p in [0,1] is the dimer density. We give upper and lower bounds for (lambda_d)(p) in terms of expressions involving (lambda_(d-1))(q). The…
Given an open set with finite perimeter $\Omega\subset \mathbb{R}^n$, we consider the space $LD_\gamma^{p}(\Omega)$, $1\leq p<\infty$, of functions with $p$th-integrable deformation tensor on $\Omega$ and with $p$ th-integrable trace value…
Erd\H{o}s asked whether every $n$-point set in Euclidean space whose $\binom{n}{2}$ pairwise distances are mutually at least $1$ apart must have diameter at least $(1+o(1))n^2$. We disprove this statement by constructing for every prime…
Let $x(n):=\alpha n^d \mod 1$ for integer $d >1$ and non-zero real $\alpha$. We show that $\{x(n)\}_{n>0}$ has Poissonian $\ell$-point correlations for almost all choices of $\alpha$ when $d$ is large (depending on $\ell$). This falls in…
The Johnson-Lindenstrauss transform is a fundamental method for dimension reduction in Euclidean spaces, that can map any dataset of $n$ points into dimension $O(\log n)$ with low distortion of their distances. This dimension bound is tight…
We use entropy numbers in combination with the polynomial method to derive a new general lower bound for the n-th minimal error in the quantum setting of information-based complexity. As an application, we improve some lower bounds on…
Let $H := \begin{pmatrix} 1 & {\mathbf R} & {\mathbf R} \\ 0 & 1 &{\mathbf R} \\ 0 & 0 & 1 \end{pmatrix}$ denote the Heisenberg group with the usual Carnot-Carath\'eodory metric $d$. It is known (since the work of Pansu and Semmes) that the…
In this paper, we introduce and develop the method of compression of points in space. We introduce the notion of the mass, the rank, the entropy, the cover and the energy of compression. We leverage this method to prove some class of…
This paper studies the minimal dimension required to embed subset memberships ($m$ elements and ${m\choose k}$ subsets of at most $k$ elements) into vector spaces, denoted as Minimal Embeddable Dimension (MED). The tight bounds of MED are…
Let $P$ be a set of $n$ points in $\mathbb{R}^d$, in general position. We remove all of them one by one, in each step erasing one vertex of the convex hull of the current remaining set. Let $g_d(P)$ denote the number of different removal…
We consider the problem of subset selection for $\ell_{p}$ subspace approximation, that is, to efficiently find a \emph{small} subset of data points such that solving the problem optimally for this subset gives a good approximation to…
The planar embedding conjecture asserts that any planar metric admits an embedding into L_1 with constant distortion. This is a well-known open problem with important algorithmic implications, and has received a lot of attention over the…
We devise a new embedding technique, which we call measured descent, based on decomposing a metric space locally, at varying speeds, according to the density of some probability measure. This provides a refined and unified framework for the…
Entropy is useful in statistical problems as a measure of irreversibility, randomness, mixing, dispersion, and number of microstates. However, there remains ambiguity over the precise mathematical formulation of entropy, generalized beyond…
Let $d \in \mathbb{N}$, $\delta \in (0, 1/2)$, and $X > 0$. Denote by $N_d(X, \delta)$ the maximum number of points in a subset of the closed Euclidean ball of radius $X$ in $\mathbb{R}^d$ such that every pairwise distance is at least…
Binary embedding is a nonlinear dimension reduction methodology where high dimensional data are embedded into the Hamming cube while preserving the structure of the original space. Specifically, for an arbitrary $N$ distinct points in…