Related papers: Mixture Models in Astronomy
In cluster analysis, it can be useful to interpret the partition built from the data in the light of external categorical variables which were not directly involved to cluster the data. An approach is proposed in the model-based clustering…
The problem of multimodal clustering arises whenever the data are gathered with several physically different sensors. Observations from different modalities are not necessarily aligned in the sense there there is no obvious way to associate…
Future cosmological galaxy surveys such as the Large Synoptic Survey Telescope (LSST) will photometrically observe very large numbers of galaxies. Without spectroscopy, the redshifts required for the analysis of these data will need to be…
Mixed data comprises both numeric and categorical features, and mixed datasets occur frequently in many domains, such as health, finance, and marketing. Clustering is often applied to mixed datasets to find structures and to group similar…
Clustering aims to divide a set of points into groups. The current paradigm assumes that the grouping is well-defined (unique) given the probability model from which the data is drawn. Yet, recent experiments have uncovered several…
We propose the use of probability models for ranked data as a useful alternative to a quantitative data analysis to investigate the outcome of bioassay experiments, when the preliminary choice of an appropriate normalization method for the…
The main tools in cosmology for comparing theoretical models with the observations of the galaxy distribution are statistical. We will review the applications of spatial statistics to the description of the large-scale structure of the…
Probabilistic finite mixture models are widely used for unsupervised clustering. These models can often be improved by adapting them to the topology of the data. For instance, in order to classify spatially adjacent data points similarly,…
We present an analysis of the integrated properties of the stellar populations in the Universidad Complutense de Madrid Survey of Halpha-selected galaxies. In this paper, the first of a series, we describe in detail the techniques developed…
The mass fraction of hot gas in clusters is a basic quantity whose level and dependence on the cluster mass and redshift are intimately linked to all cluster X-ray and SZ measures. Modeling the evolution of the gas fraction is clearly a…
We present a data-driven method to infer the redshift distribution of an arbitrary dataset based on spatial cross-correlation with a reference population and we apply it to various datasets across the electromagnetic spectrum to show its…
We review recent work that investigates the formation of stellar clusters, ranging in scale from globular clusters through open clusters to the small scale aggregates of stars observed in T associations. In all cases, recent advances in…
Traditional studies of stellar clusters in external galaxies use surface photometry and therefore focus on systems that are still bright and compact enough to be separated from the stellar background. Consequently, the latter stages of…
The abundance and mass distribution of galaxy clusters is a sensitive probe of cosmological parameters, through the sensitivity of the high-mass end of the halo mass function to $\Omega_m$ and $\sigma_8$. While galaxy cluster surveys have…
Within scientific and real life problems, classification is a typical case of extremely complex tasks in data-driven scenarios, especially if approached with traditional techniques. Machine Learning supervised and unsupervised paradigms,…
This paper is a step-by-step tutorial for fitting a mixture distribution to data. It merely assumes the reader has the background of calculus and linear algebra. Other required background is briefly reviewed before explaining the main…
We investigate the formation of clusters of galaxies in an expanding universe using a new code that regrids at a region of high density. In particular we investigate two models for the initial conditions, both with the standard CDM power…
Clustered data is ubiquitous in a variety of scientific fields. In this paper, we propose a flexible and interpretable modeling approach, called grouped heterogenous mixture modeling, for clustered data, which models cluster-wise…
Mixture models are a fundamental tool in applied statistics and machine learning for treating data taken from multiple subpopulations. The current practice for estimating the parameters of such models relies on local search heuristics…
As in other estimation scenarios, likelihood based estimation in the normal mixture set-up is highly non-robust against model misspecification and presence of outliers (apart from being an ill-posed optimization problem). A robust…