English
Related papers

Related papers: How many data clusters are in the Galaxy data set?…

200 papers

The integrated spectro-photometric properties of star clusters are subject to large cluster-to-cluster variations. They are distributed in non trivial ways around the average properties predicted by standard population synthesis models.…

Cosmology and Nongalactic Astrophysics · Physics 2009-08-20 M. Fouesneau , A. Lançon

Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a…

Methodology · Statistics 2025-10-21 Fumiya Iwashige , Tomoya Wakayama , Shonosuke Sugasawa , Shintaro Hashimoto

Star-galaxy classification is one of the most fundamental data-processing tasks in survey astronomy, and a critical starting point for the scientific exploitation of survey data. For bright sources this classification can be done with…

Instrumentation and Methods for Astrophysics · Physics 2013-07-30 Marc Henrion , Daniel J. Mortlock , David J. Hand , Axel Gandy

Cluster analysis requires many decisions: the clustering method and the implied reference model, the number of clusters and, often, several hyper-parameters and algorithms' tunings. In practice, one produces several partitions, and a final…

Machine Learning · Statistics 2023-08-14 Luca Coraggio , Pietro Coretto

Fair clustering has become a socially significant task with the advancement of machine learning technologies and the growing demand for trustworthy AI. Group fairness ensures that the proportions of each sensitive group are similar in all…

Machine Learning · Statistics 2025-06-17 Jihu Lee , Kunwoong Kim , Yongdai Kim

The paper describes clustering problems from the combinatorial viewpoint. A brief systemic survey is presented including the following: (i) basic clustering problems (e.g., classification, clustering, sorting, clustering with an order over…

Artificial Intelligence · Computer Science 2015-06-01 Mark Sh. Levin

We present the first public release of our Bayesian inference tool, Bayes-X, for the analysis of X-ray observations of galaxy clusters. We illustrate the use of Bayes-X by analysing a set of four simulated clusters at z=0.2-0.9 as they…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-17 M. Olamaie , F. Feroz , K. J. B. Grainge , M. P. Hobson , J. S. Sanders , R. D. E Saunders

Current analysis of astronomical data are confronted with the daunting task of modeling the awkward features of astronomical data, among which heteroscedastic (point-dependent) errors, intrinsic scatter, non-ignorable data collection…

Instrumentation and Methods for Astrophysics · Physics 2011-12-19 S. Andreon

In Bayesian statistics, the choice of prior distribution is often debatable, especially if prior knowledge is limited or data are scarce. In imprecise probability, sets of priors are used to accurately model and reflect prior knowledge.…

Methodology · Statistics 2016-10-25 Gero Walter , Frank P. A. Coolen

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

Methodology · Statistics 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek

We study the accuracy of Bayesian supervised method used to cluster individuals into genetically homogeneous groups on the basis of dominant or codominant molecular markers. We provide a formula relating an error criterion the number of…

Populations and Evolution · Quantitative Biology 2011-12-14 Gilles Guillot , Alexandra Carpentier-Skandalis

We present a novel framework for concomitant dimension reduction and clustering. This framework is based on a novel class of Bayesian clustering factor models. These models assume a factor model structure where the vectors of common factors…

Methodology · Statistics 2025-05-09 Hwasoo Shin , Marco A. R. Ferreira , Allison N. Tegge

Cluster analysis of biological samples using gene expression measurements is a common task which aids the discovery of heterogeneous biological sub-populations having distinct mRNA profiles. Several model-based clustering algorithms have…

Methodology · Statistics 2012-01-30 Alberto Cozzini , Ajay Jasra , Giovanni Montana

Bayesian nonparametric mixture models are widely used to cluster observations. However, one major drawback of the approach is that the estimated partition often presents unbalanced clusters' frequencies with only a few dominating clusters…

Methodology · Statistics 2026-02-03 Beatrice Franzolini , Giovanni Rebaudo

Clustering provides a common means of identifying structure in complex data, and there is renewed interest in clustering as a tool for the analysis of large data sets in many fields. A natural question is how many clusters are appropriate…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Susanne Still , William Bialek

The integrated completed likelihood (ICL) criterion has proven to be a very popular approach in model-based clustering through automatically choosing the number of clusters in a mixture model. This approach effectively maximises the…

Computation · Statistics 2015-05-26 Marco Bertoletti , Nial Friel , Riccardo Rastelli

We discuss the theoretical interpretation of observational data concerning the clustering of galaxies at high redshifts. Building on the theoretical machinery developed by Matarrese et al. (1997), we make detailed quantitative predictions…

Astrophysics · Physics 2009-10-30 Lauro Moscardini , Peter Coles , Francesco Lucchin , Sabino Matarrese

In model-based-clustering mixture models are used to group data points into clusters. A useful concept introduced for Gaussian mixtures by Malsiner Walli et al (2016) are sparse finite mixtures, where the prior distribution on the weight…

Methodology · Statistics 2018-08-23 Sylvia Frühwirth-Schnatter , Gertraud Malsiner-Walli

This article is the second in a series in which we perform an extensive comparison of various galaxy-based cluster mass estimation techniques that utilise the positions, velocities and colours of galaxies. Our aim is to quantify the…

The mixture models have become widely used in clustering, given its probabilistic framework in which its based, however, for modern databases that are characterized by their large size, these models behave disappointingly in setting out the…

Machine Learning · Statistics 2017-02-01 Abdelghafour Talibi , Boujemâa Achchab , Rafik Lasri