Related papers: Nearest Neighbor distributions: new statistical me…
We present the methodology for deriving accurate and reliable cosmological constraints from non-linear scales (<50Mpc/h) with k-th nearest neighbor (kNN) statistics. We detail our methods for choosing robust minimum scale cuts and…
Quickest change detection (QCD) is a fundamental problem in many applications. Given a sequence of measurements that exhibits two different distributions around a certain flipping point, the goal is to detect the change in distribution…
This paper considers the problem of estimating the cumulative distribution function and probability density function of a random variable using data quantized by uniform and non-uniform quantizers. A simple estimator is proposed based on…
We test the cosmological implications of studying galaxy clustering using a tomographic approach, by computing the galaxy two-point angular correlation function $\omega(\theta)$ in thin redshift shells using a spectroscopic-redshift galaxy…
The two-point correlation function of the galaxy distribution is a key cosmological observable that allows us to constrain the dynamical and geometrical state of our Universe. To measure the correlation function we need to know both the…
Nearest neighbor (NN) problem is an important scientific problem. The NN query, to find the closest one to a given query point among a set of points, is widely used in applications such as density estimation, pattern classification,…
Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning. Since clustering analysis is one of the best ways to find some clarity and structure within raw data, this paper…
We study a clustering problem where the goal is to maximize the coverage of the input points by $k$ chosen centers. Specifically, given a set of $n$ points $P \subseteq \mathbb{R}^d$, the goal is to pick $k$ centers $C \subseteq…
The statistical analysis of cosmic large-scale structure is most often based on simple two-point summary statistics, like the power spectrum or the two-point correlation function of a sample of galaxies or other types of tracers. In…
We use very large cosmological N--body simulations to obtain accurate predictions for the two-point correlations and power spectra of mass-limited samples of galaxy clusters. We consider two currently popular cold dark matter (CDM)…
This paper presents a novel perspective on correlation functions in the clustering analysis of the large-scale structure of the universe. We first recognise that pair counting in bins of radial separation is equivalent to evaluating…
We show how to simulate the clustering of rich clusters of galaxies using a technique based on the Zel'dovich approximation. This method well reproduces the spatial distribution of clusters obtainable from full N-body simulations at a…
Probabilistic k-nearest neighbour (PKNN) classification has been introduced to improve the performance of original k-nearest neighbour (KNN) classification algorithm by explicitly modelling uncertainty in the classification of each feature…
Apart from the role the clustering coefficient plays in the definition of the small-world phenomena, it also has great relevance for practical problems involving networked dynamical systems. To study the impact of the clustering coefficient…
Clustering data using prior domain knowledge, starting from a partially labeled set, has recently been widely investigated. Often referred to as semi-supervised clustering, this approach leverages labeled data to enhance clustering…
Weak-lensing searches for galaxy clusters are plagued by low completeness and purity, severely limiting their usefulness for constraining cosmological parameters with the cluster mass function. A significant fraction of `false positives'…
This thesis presents the analysis of the clustering of galaxies in the 6dF Galaxy Survey (6dFGS). At large separation scales the baryon acoustic oscillation (BAO) signal is detected which allows to make an absolute distance measurement at…
We present a simple method for evaluating the nonlinear biasing function of galaxies from a redshift survey. The nonlinear biasing is characterized by the conditional mean of the galaxy density fluctuation given the underlying mass density…
When analyzing galaxy clustering in multi-band imaging surveys, there is a trade-off between selecting the largest galaxy samples (to minimize the shot noise) and selecting samples with the best photometric redshift (photo-z) precision,…
We explore the near-infrared (NIR) $K$-band properties of galaxies within 93 galaxy clusters and groups using data from the 2MASS. We use X-ray properties of these clusters to pinpoint cluster centers and estimate cluster masses. By…