Related papers: Nearest Neighbor distributions: new statistical me…
For galaxy clustering to provide robust constraints on cosmological parameters and galaxy formation models, it is essential to make reliable estimates of the errors on clustering measurements. We present a new technique, based on a spatial…
Small- and intermediate-scale galaxy clustering can be used to establish the galaxy-halo connection to study galaxy formation and evolution and to tighten constraints on cosmological parameters. With the increasing precision of galaxy…
We propose a method for finding a cumulative distribution function (cdf) that minimizes the distance to a given cdf, while belonging to an ambiguity set constructed relative to another cdf and, possibly, incorporating soft information. Our…
Symmetric nonnegative matrix factorization (SymNMF) is a powerful tool for clustering, which typically uses the $k$-nearest neighbor ($k$-NN) method to construct similarity matrix. However, $k$-NN may mislead clustering since the neighbors…
K-Nearest Neighbours (k-NN) is a popular classification and regression algorithm, yet one of its main limitations is the difficulty in choosing the number of neighbours. We present a Bayesian algorithm to compute the posterior probability…
For a given statistic, A, the cosmic distribution function, Upsilon(VA), is the probability of measuring a value VA in a finite galaxy catalog. For statistics related to count-in-cells, such as factorial moments, F_k, the average…
The mark weighted correlation function (MCF) $W(s,\mu)$ is a computationally efficient statistical measure which can probe clustering information beyond that of the conventional 2-point statistics. In this work, we extend the traditional…
Consider a setting with multiple units (e.g., individuals, cohorts, geographic locations) and outcomes (e.g., treatments, times, items), where the goal is to learn a multivariate distribution for each unit-outcome entry, such as the…
For more than two decades, the Navarro, Frenk, and White (NFW) model has stood the test of time; it has been used to describe the distribution of mass in galaxy clusters out to their outskirts. Stacked weak lensing measurements of clusters…
Studies of disordered heterogeneous media and galaxy cosmology share a common goal: analyzing the distribution of particles at `microscales' to predict physical properties at `macroscales', whether for a liquid, composite material, or…
The number density of galaxy clusters across mass and redshift has been established as a powerful cosmological probe. Cosmological analyses with galaxy clusters traditionally employ scaling relations. However, many challenges arise from…
The 2dF Galaxy Redshift Survey has now been completed and has mapped the three-dimensional distribution, and hence clustering, of galaxies in exquisite detail over an unprecedentedly large ($\sim 10^{8} h^{-3}$ Mpc$^{3}$) volume of the…
Shape dependence of higher order correlations introduces complication in direct determination of these quantities. For this reason theoretical and observational progress has been restricted in calculating one point distribution functions…
The $k$-nearest neighbor ($k$-NN) algorithm is one of the most popular methods for nonparametric classification. However, a relevant limitation concerns the definition of the number of neighbors $k$. This parameter exerts a direct impact on…
Cluster analysis which focuses on the grouping and categorization of similar elements is widely used in various fields of research. Inspired by the phenomenon of atomic fission, a novel density-based clustering algorithm is proposed in this…
We present a comparison of major methodologies of fast generating mock halo or galaxy catalogues. The comparison is done for two-point and the three-point clustering statistics. The reference catalogues are drawn from the BigMultiDark…
Percolation analysis has long been used to quantify the connectivity of the cosmic web. Most of the previous work is based on density fields on grids. By smoothing into fields, we lose information about galaxy properties like shape or…
We present a first analysis of the clustering of SDSS galaxies using the distribution function of the sum of Fourier phases. This statistic was recently proposed by one of authors as a new method to probe phase correlations of cosmological…
Estimates of finite population cumulativedistribution functions (CDFs) and quantiles are critical forpolicy-making, resource allocation, and public health planning. For instance, federal finance agencies may require accurate estimates of…
Clustering by fast search and find of density peaks (DPC) (Since, 2014) has been proven to be a promising clustering approach that efficiently discovers the centers of clusters by finding the density peaks. The accuracy of DPC depends on…