English
Related papers

Related papers: Euclidean distance compression via deep random fea…

200 papers

This paper argues that randomized linear sketching is a natural tool for on-the-fly compression of data matrices that arise from large-scale scientific simulations and data collection. The technical contribution consists in a new algorithm…

Numerical Analysis · Computer Science 2019-02-26 Joel A. Tropp , Alp Yurtsever , Madeleine Udell , Volkan Cevher

We introduce Deep Set Linearized Optimal Transport, an algorithm designed for the efficient simultaneous embedding of point clouds into an $L^2-$space. This embedding preserves specific low-dimensional structures within the Wasserstein…

Machine Learning · Computer Science 2024-01-04 Scott Mahan , Caroline Moosmüller , Alexander Cloninger

We propose a new unsupervised anomaly detection method based on the sliced-Wasserstein distance for training data selection in machine learning approaches. Our filtering technique is interesting for decision-making pipelines deploying…

Machine Learning · Computer Science 2025-04-18 Julien Pallage , Antoine Lesage-Landry

We propose a new embedding method which is particularly well-suited for settings where the sample size greatly exceeds the ambient dimension. Our technique consists of partitioning the space into simplices and then embedding the data points…

Machine Learning · Computer Science 2020-02-07 Lee-Ad Gottlieb , Eran Kaufman , Aryeh Kontorovich , Gabriel Nivasch , Ofir Pele

Squared Wasserstein distance is a frequently used tool to measure discrepancy between probability distributions. This distance is typically computed between empirical measures of size $n$ from two underlying random samples. Unfortunately,…

Machine Learning · Statistics 2026-05-20 Peter Matthew Jacobs , Jeff M. Phillips

Although recovering an Euclidean distance matrix from noisy observations is a common problem in practice, how well this could be done remains largely unknown. To fill in this void, we study a simple distance matrix estimate based upon the…

Machine Learning · Statistics 2014-09-18 Luwan Zhang , Grace Wahba , Ming Yuan

Weak gravitational lensing, resulting from the bending of light due to the presence of matter along the line of sight, is a potent tool for exploring large-scale structures, particularly in quantifying non-Gaussianities. It stands as a…

Cosmology and Nongalactic Astrophysics · Physics 2024-06-17 Vilasini Tinnaneri Sreekanth , Sandrine Codis , Alexandre Barthelemy , Jean-Luc Starck

We present a smooth probabilistic reformulation of $\ell_0$ regularized regression that does not require Monte Carlo sampling and allows for the computation of exact gradients, facilitating rapid convergence to local optima of the best…

Machine Learning · Computer Science 2025-09-19 Lukas Silvester Barth , Paulo von Petersenn

Recent years have witnessed a tremendous growth using topological summaries, especially the persistence diagrams (encoding the so-called persistent homology) for analyzing complex shapes. Intuitively, persistent homology maps a potentially…

Computational Geometry · Computer Science 2021-04-19 Samantha Chen , Yusu Wang

For $\ell\colon \mathbb{R}^d \to [0,\infty)$ we consider the sequence of probability measures $\left(\mu_n\right)_{n \in \mathbb{N}}$, where $\mu_n$ is determined by a density that is proportional to $\exp(-n\ell)$. We allow for infinitely…

Probability · Mathematics 2023-12-11 Mareike Hasenpflug , Daniel Rudolf , Björn Sprungk

We investigate rigidity-type problems on the real line and the circle in the non-generic setting. Specifically, we consider the problem of uniquely determining the positions of $n$ distinct points $V = {v_1, \ldots, v_n}$ given a set of…

Metric Geometry · Mathematics 2024-01-30 Itai Benjamini , Elad Tzalik

Despite many applications, dimensionality reduction in the $\ell_1$-norm is much less understood than in the Euclidean norm. We give two new oblivious dimensionality reduction techniques for the $\ell_1$-norm which improve exponentially…

Data Structures and Algorithms · Computer Science 2021-08-09 Yi Li , David P. Woodruff , Taisuke Yasuda

We develop a new unsupervised symmetry learning method that starts with raw data and provides the minimal generator of an underlying Lie group of symmetries, together with a symmetry-equivariant representation of the data, which turns the…

Machine Learning · Computer Science 2025-07-08 Onur Efe , Arkadas Ozakin

In second-order optimization, a potential bottleneck can be computing the Hessian matrix of the optimized function at every iteration. Randomized sketching has emerged as a powerful technique for constructing estimates of the Hessian which…

Optimization and Control · Mathematics 2021-07-16 Michał Dereziński , Jonathan Lacotte , Mert Pilanci , Michael W. Mahoney

Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic optimal transport…

Machine Learning · Computer Science 2026-05-15 Ao Xu , Tieru Wu

The need for efficiently comparing and representing datasets with unknown alignment spans various fields, from model analysis and comparison in machine learning to trend discovery in collections of medical datasets. We use manifold learning…

Machine Learning · Statistics 2022-07-13 Tal Shnitzer , Mikhail Yurochkin , Kristjan Greenewald , Justin Solomon

This paper introduces a methodology based on Euclidean information theory to investigate local properties of secure communication over discrete memoryless wiretap channels. We formulate a constrained optimization problem that maximizes a…

Information Theory · Computer Science 2026-05-14 Emmanouil M. Athanasakos , Nicholas Kalouptsidis , Hariprasad Manjunath

In this paper we study constrained subspace approximation problem. Given a set of $n$ points $\{a_1,\ldots,a_n\}$ in $\mathbb{R}^d$, the goal of the {\em subspace approximation} problem is to find a $k$ dimensional subspace that best…

Data Structures and Algorithms · Computer Science 2025-04-30 Aditya Bhaskara , Sepideh Mahabadi , Madhusudhan Reddy Pittu , Ali Vakilian , David P. Woodruff

In compressed sensing (CS) framework, a signal is sampled below Nyquist rate, and the acquired compressed samples are generally random in nature. However, for efficient estimation of the actual signal, the sensing matrix must preserve the…

Information Theory · Computer Science 2015-07-28 V. Abrol , P. Sharma , A. K Sao

There is currently a gap in theory for point patterns that lie on the surface of objects, with researchers focusing on patterns that lie in a Euclidean space, typically planar and spatial data. Methodology for planar and spatial data thus…

Statistics Theory · Mathematics 2020-02-11 Scott Ward , Edward A. K. Cohen , Niall Adams