Related papers: Field Formulation of Parzen Data Analysis
The Diffusion Map is a nonlinear dimensionality reduction technique used to analyze high-dimensional data, with recent applications extending to datasets from the social sciences. Previous research has given little attention to how the…
The clustering of galaxy clusters is a powerful cosmological tool, which can help to break degeneracies between parameters when combined with other cosmological observables. We aim to demonstrate its potential in constraining cosmological…
Spectral clustering is a powerful method for finding structure in a dataset through the eigenvectors of a similarity matrix. It often outperforms traditional clustering algorithms such as $k$-means when the structure of the individual…
We consider the signed density of the extremal points of (two-dimensional) scalar fields with a Gaussian distribution. We assign a positive unit charge to the maxima and minima of the function and a negative one to its saddles. At first, we…
In graph motivated learning, label propagation largely depends on data affinity represented as edges between connected data points. The affinity assignment implicitly assumes even distribution of data on the manifold. This assumption may…
We explore the implications of a single observer's viewpoint on measurements of galaxy clustering statistics. We focus on the Bardeen potentials, which imprint characteristic scale-dependent signatures in the observed galaxy power spectrum.…
Clustering is an unsupervised technique for grouping data points by similarity. While explainability methods exist for supervised machine learning, they are not directly applicable to clustering, making it challenging to understand cluster…
Likelihood fitting to two-point clustering statistics made from galaxy surveys usually assumes a multivariate normal distribution for the measurements, with justification based on the central limit theorem given the large number of…
Dictionary learning and sparse coding have been widely studied as mechanisms for unsupervised feature learning. Unsupervised learning could bring enormous benefit to the processing of hyperspectral images and to other remote sensing data…
The number of modes in a probability density function is representative of the complexity of a model and can also be viewed as the number of subpopulations. Despite its relevance, there has been limited research in this area. A novel…
We present in this article an analysis of some of the properties of the density field realized in numerical simulations for power-law initial power-spectra in the case of a critical density universe. We compare our numerical results in the…
We prove finite-sample concentration and anti-concentration bounds for dimension estimation using Gaussian kernel sums. Our bounds provide explicit dependence on sample size, bandwidth, and local geometric and distributional parameters,…
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of $m$ points in $n$ dimensions, $n,m \rightarrow \infty$ and $\alpha = m/n$ stays finite. Using exact but non-rigorous methods…
The eigenvalue probability density function (PDF) for the Gaussian unitary ensemble has a well known analogy with the Boltzmann factor for a classical log-gas with pair potential $- \log | x - y|$, confined by a one-body harmonic potential.…
A general framework for dealing with both linear regression and clustering problems is described. It includes Gaussian clusterwise linear regression analysis with random covariates and cluster analysis via Gaussian mixture models with…
Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the…
This work further develops the calculation of QED effects in a finite Gaussian basis. We focus on the non-linear ${\alpha}(Z{\alpha})^{n\ge 3}$ contribution to the vacuum polarization density, computing the energy shift of 1s$_{1/2}$ states…
This paper focuses on obtaining clustering information about a distribution from its i.i.d. samples. We develop theoretical results to understand and use clustering information contained in the eigenvectors of data adjacency matrices based…
Weak gravitational lensing surveys have the potential to directly probe mass density fluctuation in the universe. Recent studies have shown that it is possible to model the statistics of the convergence field at small angular scales by…
One of the fundamental problems in machine learning is the estimation of a probability distribution from data. Many techniques have been proposed to study the structure of data, most often building around the assumption that observations…