Related papers: Nearest Neighbor distributions: new statistical me…
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, density estimation itself is also a challenging problem, especially the determination of the kernel bandwidth. A large bandwidth…
We present here a new algorithm for the fast computation of N-point correlation functions in large astronomical data sets. The algorithm is based on kdtrees which are decorated with cached sufficient statistics thus allowing for orders of…
KNN has the reputation to be the word simplest but efficient supervised learning algorithm used for either classification or regression. KNN prediction efficiency highly depends on the size of its training data but when this training data…
When data is of an extraordinarily large size or physically stored in different locations, the distributed nearest neighbor (NN) classifier is an attractive tool for classification. We propose a novel distributed adaptive NN classifier for…
We present a new algorithm for efficiently computing the $N$-point correlation functions (NPCFs) of a 3D density field for arbitrary $N$. This can be applied both to a discrete spectroscopic galaxy survey and a continuous field. By…
The "spectral correlation function" analysis we introduce in this paper is a new tool for analyzing spectral-line data cubes. Our initial tests, carried out on a suite of observed and simulated data cubes, indicate that the spectral…
In this work, we discuss a general class of the estimators for the cumulative distribution function (CDF) based on judgment post stratification (JPS) sampling scheme which includes both empirical and kernel distribution functions.…
Motivated by the needs of estimating the proximity clustering with partial distance measurements from vantage points or landmarks for remote networked systems, we show that the proximity clustering problem can be effectively formulated as…
Using a discrete wavelet based space-scale decomposition (SSD), the spectrum of the skewness and kurtosis is developed to describe the non-Gaussian signatures in cosmologically interesting samples. Because the basis of the discrete wavelet…
We have simulated the growth of structure in two 100 Mpc boxes for LCDM and SCDM universes. These N-body/SPH simulations include a gaseous component which is able to cool radiatively. A fraction of the gas cools into cold dense objects…
A quantile is defined as a value below which random draws from a given distribution falls with a given probability. In a centralized setting where the cumulative distribution function (CDF) is unknown, the empirical CDF (ECDF) can be used…
We introduce a novel approach for studying random k-coverage, using Morse theory for the k-nearest neighbor (k-NN) distance function. We prove a sharp phase transition for the number of critical points of the k-NN distance function, from…
We study the potential of weak lensing surveys to detect clusters of galaxies, using a fast Particle Mesh cosmological N-body simulation algorithm specifically tailored to investigate the statistics of these mass-selected clusters. In…
In the era of big data, k-means clustering has been widely adopted as a basic processing tool in various contexts. However, its computational cost could be prohibitively high as the data size and the cluster number are large. It is well…
We present a data-driven method to infer the redshift distribution of an arbitrary dataset based on spatial cross-correlation with a reference population and we apply it to various datasets across the electromagnetic spectrum to show its…
We present new predictions for the galaxy three-point correlation function (3PCF) using high-resolution dissipationless cosmological simulations of a flat LCDM Universe which resolve galaxy-size halos and subhalos. We create realistic mock…
In this paper a relative number density parameter, called the neighborhood function, is introduced so that the crowded nature of the neighborhood of individual sources can be described. With this parameter one can determine the probability…
Two-point correlation functions (2PCF) are widely used to characterize how points cluster in space. In this work, we study the problem of measuring the 2PCF over a large set of points, restricted to a subset satisfying a property of…
The k-Nearest Neighbor (k-NN) classification algorithm is one of the most widely-used lazy classifiers because of its simplicity and ease of implementation. It is considered to be an effective classifier and has many applications. However,…
The cumulative distribution network (CDN) is a recently developed class of probabilistic graphical models (PGMs) permitting a copula factorization, in which the CDF, rather than the density, is factored. Despite there being much recent…