Related papers: Nearest Neighbor distributions: new statistical me…
We study the following distribution clustering problem: Given a hidden partition of $k$ distributions into two groups, such that the distributions within each group are the same, and the two distributions associated with the two clusters…
The two-point correlation function (2PCF) is the most widely used tool for quantifying the spatial distribution of galaxies. Since the distribution of galaxies is determined by galaxy formation physics as well as the underlying cosmology,…
The simple one-parameter nearest neighbor-spacing distribution (NNSD) is suggested for statistical analysis of nuclear spectra. This distribution is derived within the Wigner-Dyson approach in the linear approximation for the level…
We have extracted over 400 clusters, covering more than 2 decades in mass, from three simulations of the TCDM cosmology. This represents the largest, uniform catalogue of simulated clusters ever produced. The clusters exhibit a wide variety…
In two recent papers, we developed a powerful technique to link the distribution of galaxies to that of dark matter haloes by considering halo occupation numbers as function of galaxy luminosity and type. In this paper we use these…
In the $k$-nearest neighborhood model ($k$-NN), we are given a set of points $P$, and we shall answer queries $q$ by returning the $k$ nearest neighbors of $q$ in $P$ according to some metric. This concept is crucial in many areas of data…
The two point correlation function (2PCF) is a powerful statistical tool to measure galaxy clustering. Although 2PCF has also been used to study the clustering of stars on parsec and sub-parsec scales, its physical implication is not clear…
We present an implicit likelihood approach to quantifying cosmological information over discrete catalogue data, assembled as graphs. To do so, we explore cosmological parameter constraints using mock dark matter halo catalogues. We employ…
Learning the multivariate distribution of data is a core challenge in statistics and machine learning. Traditional methods aim for the probability density function (PDF) and are limited by the curse of dimensionality. Modern neural methods…
We develop a diagrammatic technique to represent the multi-point cumulative probability density function (CPDF) of mass fluctuations in terms of the statistical properties of individual collapsed objects and relate this to other statistical…
Combining galaxy clustering information from regions of different environmental densities can help break cosmological parameter degeneracies and access non-Gaussian information from the density field that is not readily captured by the…
We discuss a graph-based approach for testing spatial point patterns. This approach falls under the category of data-random graphs, which have been introduced and used for statistical pattern recognition in recent years. Our goal is to test…
In the context of count-in-cells statistics, the joint probability distribution of the density in two concentric spherical shells is predicted from first first principle for sigmas of the order of one. The agreement with simulation is found…
Density based spatial clustering of points in $\mathbb{R}^n$ has a myriad of applications in a variety of industries. We generalise this problem to the density based clustering of lines in high-dimensional spaces, keeping in mind there…
In this letter we explore the suggestion of Quashnock and Lamb (1993) that nearest neighbor correlations among gamma ray burst positions indicate the possibility of burst repetitions within various burst sub-classes. With the aid of Monte…
Spatial interaction between two or more classes or species has important implications in various fields and causes multivariate patterns such as segregation or association. Segregation occurs when members of a class or species are more…
We introduce methods which allow observed galaxy clustering to be used together with observed luminosity or stellar mass functions to constrain the physics of galaxy formation. We show how the projected two-point correlation function of…
This paper presents how to perform minimax optimal classification, regression, and density estimation based on fixed-$k$ nearest neighbor (NN) searches. We consider a distributed learning scenario, in which a massive dataset is split into…
Graph measures that express closeness or distance between nodes can be employed for graph nodes clustering using metric clustering algorithms. There are numerous measures applicable to this task, and which one performs better is an open…
A novel nonparametric clustering algorithm is proposed using the interpoint distances between the members of the data to reveal the inherent clustering structure existing in the given set of data, where we apply the classical nonparametric…