English
Related papers

Related papers: How many data clusters are in the Galaxy data set?…

200 papers

In computational biology, gene expression datasets are characterized by very few individual samples compared to a large number of measurements per sample. Thus, it is appealing to merge these datasets in order to increase the number of…

Methodology · Statistics 2011-08-18 Meili Baragatti

The measurement of the efficiency of an event selection is always an important part of the analysis of experimental data. The statistical techniques which are needed to determine the efficiency and its uncertainty are reviewed. Frequentist…

Data Analysis, Statistics and Probability · Physics 2012-08-28 Diego Casadei

We consider the Bayesian mixture of finite mixtures (MFMs) and Dirichlet process mixture (DPM) models for clustering. Recent asymptotic theory has established that DPMs overestimate the number of clusters for large samples and that…

Clustering task of mixed data is a challenging problem. In a probabilistic framework, the main difficulty is due to a shortage of conventional distributions for such data. In this paper, we propose to achieve the mixed data clustering with…

Methodology · Statistics 2015-10-01 Matthieu Marbac , Christophe Biernacki , Vincent Vandewalle

We develop a Bayesian model for globular clusters composed of multiple stellar populations, extending earlier statistical models for open clusters composed of simple (single) stellar populations (vanDyk et al. 2009, Stein et al. 2013).…

Solar and Stellar Astrophysics · Physics 2016-07-27 D. C. Stenning , R. Wagner-Kaiser , E. Robinson , D. A. van Dyk , T. von Hippel , A. Sarajedini , N. Stein

Low-rank matrix estimation from incomplete measurements recently received increased attention due to the emergence of several challenging applications, such as recommender systems; see in particular the famous Netflix challenge. While the…

Machine Learning · Statistics 2014-10-23 Pierre Alquier , Vincent Cottet , Nicolas Chopin , Judith Rousseau

We study methods for reconstructing Bayesian uncertainties on dynamical mass estimates of galaxy clusters using convolutional neural networks (CNNs). We discuss the statistical background of approximate Bayesian neural networks and…

Cosmology and Nongalactic Astrophysics · Physics 2021-03-16 Matthew Ho , Arya Farahi , Markus Michael Rau , Hy Trac

The presence or absence of star clusters in galaxies, and the properties of star cluster populations compared to their host galaxy properties, are important observables for validating models of cluster formation, galaxy formation, and…

Astrophysics of Galaxies · Physics 2023-08-09 Samantha C. Berek , Marta Reina-Campos , Gwendolyn Eadie , Alison Sills

In reliability engineering, data about failure events is often scarce. To arrive at meaningful estimates for the reliability of a system, it is therefore often necessary to also include expert information in the analysis, which is…

Methodology · Statistics 2016-10-25 Gero Walter , Frank P. A. Coolen

The detection of galaxy clusters in present and future surveys enables measuring mass-to-light ratios, clustering properties or galaxy cluster abundances and therefore, constraining cosmological parameters. We present a new technique for…

Cosmology and Nongalactic Astrophysics · Physics 2010-11-17 Begoña Ascaso , David Wittman , Narciso Benítez , the DLS collaboration

Global data association is an essential prerequisite for robot operation in environments seen at different times or by different robots. Repetitive or symmetric data creates significant challenges for existing methods, which typically rely…

Robotics · Computer Science 2025-09-22 Yixuan Jia , Mason B. Peterson , Qingyuan Li , Yulun Tian , Jonathan P. How

The Bayesian approach to clustering is often appreciated for its ability to provide uncertainty in the partition structure. However, summarizing the posterior distribution over the clustering structure can be challenging, due the discrete,…

Computation · Statistics 2026-01-26 Cecilia Balocchi , Sara Wade

Dirichlet process mixtures are flexible non-parametric models, particularly suited to density estimation and probabilistic clustering. In this work we study the posterior distribution induced by Dirichlet process mixtures as the sample size…

Statistics Theory · Mathematics 2022-11-29 Filippo Ascolani , Antonio Lijoi , Giovanni Rebaudo , Giacomo Zanella

Bayesian models offer great flexibility for clustering applications---Bayesian nonparametrics can be used for modeling infinite mixtures, and hierarchical Bayesian models can be utilized for sharing clusters across multiple data sets. For…

Machine Learning · Computer Science 2012-06-15 Brian Kulis , Michael I. Jordan

Lensing by galaxy clusters is a versatile probe of cosmology and extragalactic astrophysics, but the accuracy of some of its predictions is limited by the simplified models adopted to reduce the (otherwise untractable) number of degrees of…

Cosmology and Nongalactic Astrophysics · Physics 2021-04-28 Pietro Bergamini , Adriano Agnello , Gabriel Bartosch Caminha

The detection of galaxy clusters in present and future surveys enables measuring mass-to-light ratios, clustering properties, galaxy cluster abundances and therefore, constraining cosmological parameters. We present a new technique for…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-03 Begoña Ascaso , David M. Wittman , Narciso Benítez

Non-Gaussian mixture models are gaining increasing attention for mixture model-based clustering particularly when dealing with data that exhibit features such as skewness and heavy tails. Here, such a mixture distribution is presented,…

Computation · Statistics 2020-05-07 Yuan Fang , Dimitris Karlis , Sanjeena Subedi

Research on cluster analysis for categorical data continues to develop, with new clustering algorithms being proposed. However, in this context, the determination of the number of clusters is rarely addressed. In this paper, we propose a…

Methodology · Statistics 2014-09-29 Cláudia Silvestre , Margarida G. M. S. Cardoso , Mário A. T. Figueiredo

Determining the number of clusters is a fundamental issue in data clustering. Several algorithms have been proposed, including centroid-based algorithms using the Euclidean distance and model-based algorithms using a mixture of probability…

Machine Learning · Computer Science 2024-07-30 Ryosuke Motegi , Yoichi Seki

A central problem in analyzing networks is partitioning them into modules or communities. One of the best tools for this is the stochastic block model, which clusters vertices into blocks with statistically homogeneous pattern of links.…

Machine Learning · Statistics 2016-05-24 Xiaoran Yan
‹ Prev 1 3 4 5 6 7 10 Next ›