Related papers: Interpreting the Distance Correlation Results for …
We have used a Galactic sample of OB stars and associations to test the performance of an automatic grouping algorithm designed to identify extragalactic OB associations. The algorithm identifies the known Galactic OB associations correctly…
A geometrical pattern is a set of points with all pairwise distances (or, more generally, relative distances) specified. Finding matches to such patterns has applications to spatial data in seismic, astronomical, and transportation…
Time series clustering is an unsupervised learning method for classifying time series data into groups with similar behavior. It is used in applications such as healthcare, finance, economics, energy, and climate science. Several time…
Cross-correlations between datasets are used in many different contexts in cosmological analyses. Recently, $k$-Nearest Neighbor Cumulative Distribution Functions ($k{\rm NN}$-${\rm CDF}$) were shown to be sensitive probes of cosmological…
The Pearson distance between a pair of random variables $X,Y$ with correlation $\rho_{xy}$, namely, 1-$\rho_{xy}$, has gained widespread use, particularly for clustering, in areas such as gene expression analysis, brain imaging and cyber…
We consider two parametrized random digraph families, namely, proportional-edge and central similarity proximity catch digraphs (PCDs) and compare the performance of these two PCD families in testing spatial point patterns. These PCD…
We detect angular galaxy-QSO cross-correlations between the APM Galaxy Catalogue and a preliminary release (consisting of roughly half of the anticipated final catalogue) of the Hamburg-ESO Catalogue of Bright QSOs as a function of source…
Studies of the distribution and evolution of galaxies are of fundamental importance to modern cosmology; these studies, however, are hampered by the complexity of the competing effects of spectral and density evolution. Constructing a…
Longslit spectroscopy is entering an era of increased spatial and spectral resolution and increased sample size. Improved instruments reveal complex velocity structure that cannot be described with a one-dimensional rotation curve, yet…
We investigate the use of the cross-correlation between galaxies and galaxy groups to measure redshift-space distortions (RSD) and thus probe the growth rate of cosmological structure. This is compared to the classical approach based on…
We use high-resolution N-body simulations to develop a new, flexible, empirical approach for measuring the growth rate from redshift-space distortions (RSD) in the 2-point galaxy correlation function. We quantify the systematic error in…
A measure of distance between two clusterings has important applications, including clustering validation and ensemble clustering. Generally, such distance measure provides navigation through the space of possible clusterings. Mostly used…
We present applications of statistical data analysis methods from both bi- and multivariate statistics to find suitable sets of neutron star features that can be leveraged for accurate and EoS independent -- or universal -- relations. To…
A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of…
Measuring the distance between data points is fundamental to many statistical techniques, such as dimension reduction or clustering algorithms. However, improvements in data collection technologies has led to a growing versatility of…
The evolution of galaxy cluster counts is a powerful probe of several fundamental cosmological parameters. A number of recent studies using this probe have claimed tension with the cosmology preferred by the analysis of the Planck primary…
We study the problem of applying spectral clustering to cluster multi-scale data, which is data whose clusters are of various sizes and densities. Traditional spectral clustering techniques discover clusters by processing a similarity…
A typical galaxy survey geometry results in galaxy pairs of different separation and angle to the line-of-sight having different distributions in redshift and consequently a different effective redshift. However, clustering measurements are…
The angular correlation is a method for measuring the distribution of structure in the Universe, through the statistical properties of the angular distribution of galaxies on the sky. We measure the angular correlation of galaxies from the…
In machine learning, observation features are measured in a metric space to obtain their distance function for optimization. Given similar features that are statistically sufficient as a population, a statistical distance between two…