Related papers: Linear Dimension Reduction Approximately Preservin…
It is proposed that the propagation of light in disordered photonic lattices can be harnessed as a random projection that preserves distances between a set of projected vectors. This mapping is enabled by the complex evolution matrix of a…
Real-world data usually have high dimensionality and it is important to mitigate the curse of dimensionality. High-dimensional data are usually in a coherent structure and make the data in relatively small true degrees of freedom. There are…
Statistical distance measures have found wide applicability in information retrieval tasks that typically involve high dimensional datasets. In order to reduce the storage space and ensure efficient performance of queries, dimensionality…
It is shown that the correlation functions of the random variables $\det(\lambda - X)$, in which $X$ is a real symmetric $ N\times N$ random matrix, exhibit universal local statistics in the large $N$ limit. The derivation relies on an…
A basic problem in machine learning is to find a mapping $f$ from a low dimensional latent space $\mathcal{Y}$ to a high dimensional observation space $\mathcal{X}$. Modern tools such as deep neural networks are capable to represent general…
Objective functions that optimize deep neural networks play a vital role in creating an enhanced feature representation of the input data. Although cross-entropy-based loss formulations have been extensively used in a variety of supervised…
We introduce the concept of ODD ('$\mathbf{O}$rthogonally $\mathbf{D}$egenerating on a $\mathbf{D}$ivisor') Riemannian metrics on real analytic manifolds $M$. These semipositive symmetric $2$-tensors may degenerate on a finite collection of…
When performing classification tasks, raw high dimensional features often contain redundant information, and lead to increased computational complexity and overfitting. In this paper, we assume the data samples lie on a single underlying…
We study the following basic machine learning task: Given a fixed set of $d$-dimensional input points for a linear regression problem, we wish to predict a hidden response value for each of the points. We can only afford to attain the…
We initiate the rigorous study of classification in semimetric spaces, which are point sets with a distance function that is non-negative and symmetric, but need not satisfy the triangle inequality. For metric spaces, the doubling dimension…
Data augmentation is one of the most popular techniques for improving the robustness of neural networks. In addition to directly training the model with original samples and augmented samples, a torrent of methods regularizing the distance…
We describe a framework in which is possible to develop and implement algorithms for the approximation of invariant measures of dynamical systems with a given bound on the error of the approximation. Our approach is based on a general…
As network data has become ubiquitous in the sciences, there has been growing interest in network models whose structure is driven by latent node-level variables in a (typically low-dimensional) latent geometric space. These "latent…
This work studies an explicit embedding of the set of probability measures into a Hilbert space, defined using optimal transport maps from a reference probability density. This embedding linearizes to some extent the 2-Wasserstein space,…
Scientists and engineers rely on accurate mathematical models to quantify the objects of their studies, which are often high-dimensional. Unfortunately, high-dimensional models are inherently difficult, i.e. when observations are sparse or…
The Whitney embedding theorem gives an upper bound on the smallest embedding dimension of a manifold. If a data set lies on a manifold, a random projection into this reduced dimension will retain the manifold structure. Here we present an…
While there is extensive literature on approximation, deterministic as well as random, of general convex bodies $K$ in the symmetric difference metric, or other metrics arising from intrinsic volumes, very little is known for corresponding…
Low-dimensional embedding, manifold learning, clustering, classification, and anomaly detection are among the most important problems in machine learning. The existing methods usually consider the case when each instance has a fixed,…
A comprehensive approach to Sobolev-type embeddings, involving arbitrary rearrangement- invariant norms on the entire Euclidean space R^n, is offered. In particular, the optimal target space in any such embedding is exhibited. Crucial in…
We demonstrate $k+1$-term arithmetic progressions in certain subsets of the real line whose "higher-order Fourier dimension" is sufficiently close to 1. This Fourier dimension, introduced in previous work, is a higher-order (in the sense of…