Related papers: Nearest Neighbor distributions: new statistical me…
This paper is the first in a set that analyses the covariance matrices of clustering statistics obtained from several approximate methods for gravitational structure formation. We focus here on the covariance matrices of anisotropic…
We study the sample complexity of learning a uniform approximation of an $n$-dimensional cumulative distribution function (CDF) within an error $\epsilon > 0$, when observations are restricted to a minimal one-bit feedback. This serves as a…
(Abridged) This is the first of a series of papers in which we derive simultaneous constraints on cosmological parameters and X-ray scaling relations using observations of the growth of massive, X-ray flux-selected galaxy clusters. Our data…
We investigate sample-based learning of conditional distributions on multi-dimensional unit boxes, allowing for different dimensions of the feature and target spaces. Our approach involves clustering data near varying query points in the…
The upcoming XMM Large Scale Structure Survey (XMM-LSS) will ultimately provide a unique mapping of the distribution of X-ray sources in a contiguous 64 sq. deg. region. In particular, it will provide the 3-dimensional location of about 900…
We explore a signature of phase correlations in Fourier modes of dark matter density fields induced by nonlinear gravitational clustering. We compute the distribution function of the phase sum of the Fourier modes,…
Thanks to the recent availability of large surveys, there has been renewed interest in third-order correlation statistics. Measures of third-order clustering are sensitive to the structure of filaments and voids in the universe and are…
Clustering is a fundamental analysis tool aiming at classifying data points into groups based on their similarity or distance. It has found successful applications in all natural and social sciences, including biology, physics, economics,…
Spatial clustering is a crucial field, finding universal use across criminology, pathology, and urban planning. However, most spatial clustering algorithms cannot pull information from nearby nodes and suffer performance drops when dealing…
This paper proposes a new probabilistic classification algorithm using a Markov random field approach. The joint distribution of class labels is explicitly modelled using the distances between feature vectors. Intuitively, a class label…
(abridged) The clustering properties of galaxies belonging to different luminosity ranges or having different morphological types are different. These characteristics or `marks' permit to understand the galaxy catalogs that carry all this…
The alignment of clusters of galaxies with their nearest neighbours and between clusters within a supercluster is investigated using simulations of 512^{3} dark matter particles for \LambdaCDM and \tauCDM cosmological models. Strongly…
The k-Nearest Neighbor (kNN) classification approach is conceptually simple - yet widely applied since it often performs well in practical applications. However, using a global constant k does not always provide an optimal solution, e.g.,…
Context: Two-point correlation functions are used throughout cosmology as a measure for the statistics of random fields. When used in Bayesian parameter estimation, their likelihood function is usually replaced by a Gaussian approximation.…
Spatial clustering detection has a variety of applications in diverse fields, including identifying infectious disease outbreaks, assessing land use patterns, pinpointing crime hotspots, and identifying clusters of neurons in brain imaging…
We consider a power-constrained sensor network, consisting of multiple sensor nodes and a fusion center (FC), that is deployed for the purpose of estimating a common random parameter of interest. In contrast to the distributed framework,…
The k Nearest Neighbors (kNN) method has received much attention in the past decades, where some theoretical bounds on its performance were identified and where practical optimizations were proposed for making it work fairly well in high…
Statistical errors in ground state observables and single-particle properties of spherical even-even nuclei and their propagation to the limits of nuclear landscape have been investigated in covariant density functional theory (CDFT) for…
The two-point correlation function (2pcf) is the key statistic in structure formation; it measures the clustering of galaxies or other density field tracers. Estimators of the 2pcf, including the standard Landy-Szalay (LS) estimator,…
The most widely used internal measure for clustering evaluation is the silhouette coefficient, whose naive computation requires a quadratic number of distance calculations, which is clearly unfeasible for massive datasets. Surprisingly,…