Related papers: Geometric Interpretations of the $k$-Nearest Neigh…
The use of summary statistics beyond the two-point correlation function to analyze the non-Gaussian clustering on small scales is an active field of research in cosmology. In this paper, we explore a set of new summary statistics -- the…
Cross-correlations between datasets are used in many different contexts in cosmological analyses. Recently, $k$-Nearest Neighbor Cumulative Distribution Functions ($k{\rm NN}$-${\rm CDF}$) were shown to be sensitive probes of cosmological…
Distances to the $k$-nearest-neighbor ($k$NN) data points from volume-filling query points are a sensitive probe of spatial clustering. Here we present the first application of $k$NN summary statistics to observational clustering…
Searches for primordial non-Gaussianity in cosmological perturbations are a key means of revealing novel primordial physics. However, robustly extracting signatures of primordial non-Gaussianity from non-linear scales of the late-time…
This letter characterizes the statistics of the contact distance and the nearest neighbor (NN) distance for binomial point processes (BPP) spatially-distributed on spherical surfaces. We consider a setup of $n$ concentric spheres, with each…
In this letter, we derive the CDF (cumulative distribution function) of $k$th contact distance (CD) and nearest neighbor distance (NND) of the $n$-dimensional ($n$-D) Mat\'ern cluster process (MCP). We present a new approach based on the…
Understanding how atmospheric molecular clusters form and grow is key to resolving one of the biggest uncertainties in climate modelling: the formation of new aerosol particles. While quantum chemistry offers accurate insights into these…
Widefield surveys of the sky probe many clustered scalar fields -- such as galaxy counts, lensing potential, gas pressure, etc. -- that are sensitive to different cosmological and astrophysical processes. Our ability to constrain such…
In astronomy and cosmology, significant effort is devoted to characterizing and understanding spatial cross-correlations between points - e.g. galaxy positions, high energy neutrino arrival directions, X-ray and AGN sources, and continuous…
Beyond standard summary statistics are necessary to summarize the rich information on non-linear scales in the era of precision galaxy clustering measurements. For the first time, we introduce the 2D k-th nearest neighbor (kNN) statistics…
We investigate the application of Hybrid Effective Field Theory (HEFT) -- which combines a Lagrangian bias expansion with subsequent particle dynamics from $N$-body simulations -- to the modeling of $k$-Nearest Neighbor Cumulative…
Extracting cosmological parameters from galaxy/halo catalogues with sub-percent level accuracy is an important aspect of modern cosmology, especially in view of ongoing and upcoming surveys such as Euclid, DESI, and LSST. While traditional…
The dependence of galaxy clustering on local density provides an effective method for extracting non-Gaussian information from galaxy surveys. The two-point correlation function (2PCF) provides a complete statistical description of a…
Accurate modelling of redshift-space distortions (RSD) is challenging in the non-linear regime for two-point statistics e.g. the two-point correlation function (2PCF). We take a different perspective to split the galaxy density field…
Nearest neighbor search has found numerous applications in machine learning, data mining and massive data processing systems. The past few years have witnessed the popularity of the graph-based nearest neighbor search paradigm because of…
We investigate whether neighbor-density-weighted marked correlation functions (MCFs) can extract cosmological information beyond the standard redshift-space two-point correlation function (2PCF). Using the Kun suite of 129 $w_0w_a$CDM$+\sum…
Many statistical methods have been proposed in the last years for analyzing the spatial distribution of galaxies. Very few of them, however, can handle properly the border effects of complex observational sample volumes. In this paper, we…
We study the sample complexity of learning a uniform approximation of an $n$-dimensional cumulative distribution function (CDF) within an error $\epsilon > 0$, when observations are restricted to a minimal one-bit feedback. This serves as a…
Nearest neighbor (k-NN) graphs are widely used in machine learning and data mining applications, and our aim is to better understand what they reveal about the cluster structure of the unknown underlying distribution of points. Moreover, is…
In the era of big data, k-means clustering has been widely adopted as a basic processing tool in various contexts. However, its computational cost could be prohibitively high as the data size and the cluster number are large. It is well…