Related papers: Transformationally decoupling clustering and trace…
Many fundamental statistical methods have become critical tools for scientific data analysis yet do not scale tractably to modern large datasets. This paper will describe very recent algorithms based on computational geometry which have…
The growth rate of large-scale structure provides a powerful consistency test of the standard cosmological model and a probe of possible deviations from general relativity. We use a Fisher analysis to forecast constraints on the growth rate…
'Big' high-dimensional data are commonly analyzed in low-dimensions, after performing a dimensionality-reduction step that inherently distorts the data structure. For the same purpose, clustering methods are also often used. These methods…
The determination of cluster centers generally depends on the scale that we use to analyze the data to be clustered. Inappropriate scale usually leads to unreasonable cluster centers and thus unreasonable results. In this study, we first…
This paper presents an analysis of properties of two hybrid discretization methods for Gaussian derivatives, based on convolutions with either the normalized sampled Gaussian kernel or the integrated Gaussian kernel followed by central…
The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…
We present a comparative study of the accuracy and precision of correlation function methods and full-field inference in cosmological data analysis. To do so, we examine a Bayesian hierarchical model that predicts log-normal fields and…
The greatest challenge in the interpretation of galaxy clustering data from any surveys is galaxy bias. Using a simple Fisher matrix analysis, we show that the bispectrum provides an excellent determination of linear and non-linear bias…
A blurring algorithm with linear time complexity can reduce the small-scale content of data observed at scattered locations in a spatially extended domain of arbitrary dimension. The method works by forming a Gaussian interpolant of the…
We consider the shape of the posterior distribution to be used when fitting cosmological models to power spectra measured from galaxy surveys. At very large scales, Gaussian posterior distributions in the power do not approximate the…
Determining accurate redshift distributions for very large samples of objects has become increasingly important in cosmology. We investigate the impact of extending cross-correlation based redshift distribution recovery methods to include…
We propose a Fourier-based approach for optimization of several clustering algorithms. Mathematically, clusters data can be described by a density function represented by the Dirac mixture distribution. The density function can be smoothed…
In the peaks approach, the formation sites of observable structures in the Universe are identified as peaks in the matter density field. The statistical properties of the clustering of peaks are particularly important in this respect. In…
This paper develops an in-depth treatment concerning the problem of approximating the Gaussian smoothing and Gaussian derivative computations in scale-space theory for application on discrete data. With close connections to previous…
This review presents a comprehensive overview of galaxy bias, that is, the statistical relation between the distribution of galaxies and matter. We focus on large scales where cosmic density fields are quasi-linear. On these scales, the…
Skew-spectra allow us to extract non-Gaussian information by taking the square of a map and finding the power spectrum of this new map with the original map. This allows us to use much of the infrastructure of power spectra and avoid the…
An original graph clustering approach to efficient localization of error covariances is proposed within an ensemble-variational data assimilation framework. Here the localization term is very generic and refers to the idea of breaking up a…
We show how the cosmological constant can be estimated from redshift surveys at different redshifts, using maximum-likelihood techniques. The apparent redshift-space clustering on large scales (\simgt 20 \himpc) are affected in the radial…
Spatial fields in the Earth and environmental sciences are often available at multiple scales or resolutions. While coarse-scale data (e.g., from global circulation models) are often abundant, they lack the local detail provided by…
Because of its mathematical tractability, the Gaussian mixture model holds a special place in the literature for clustering and classification. For all its benefits, however, the Gaussian mixture model poses problems when the data is skewed…