English
Related papers

Related papers: Replica analysis of Bayesian data clustering

200 papers

Cluster sampling is common in survey practice, and the corresponding inference has been predominantly design-based. We develop a Bayesian framework for cluster sampling and account for the design effect in the outcome modeling. We consider…

Methodology · Statistics 2020-06-24 Susanna Makela , Yajuan Si , Andrew Gelman

This manuscript goes through the fundamental connections between statistical mechanics and estimation theory by focusing on the particular problem of compressive sensing. We first show that the asymptotic analysis of a sparse recovery…

Information Theory · Computer Science 2022-12-21 Ali Bereyhi , Ralf R. Müller , Hermann Schulz-Baldes

Subspace clustering is the classical problem of clustering a collection of data samples that approximately lie around several low-dimensional subspaces. The current state-of-the-art approaches for this problem are based on the…

Machine Learning · Computer Science 2023-01-26 Maryam Abdolali , Nicolas Gillis

In this paper, we propose a data based transformation for infinite-dimensional Gaussian processes and derive its limit theorem. For a classification problem, this transformation induces complete separation among the associated Gaussian…

Statistics Theory · Mathematics 2022-03-25 Juan A. Cuesta-Albertos , Subhajit Dutta

Influence of surrounding matter on the properties of clusters is considered by an approach combining the methods of statistical and quantum mechanics. A cluster is treated as a bound N-particle system and surrounding matter as thermostat.…

Statistical Mechanics · Physics 2015-06-25 V. I. Yukalov , E. P. Yukalova

We propose a simple and efficient clustering method for high-dimensional data with a large number of clusters. Our algorithm achieves high-performance by evaluating distances of datapoints with a subset of the cluster centres. Our…

Machine Learning · Computer Science 2022-03-30 Georgios Exarchakis , Omar Oubari , Gregor Lenz

Replication studies are essential for assessing the credibility of claims from original studies. A critical aspect of designing replication studies is determining their sample size; a too small sample size may lead to inconclusive studies…

Methodology · Statistics 2023-08-14 Samuel Pawel , Guido Consonni , Leonhard Held

If the prior probability distributions of all possible hypothetical true means and all possible observed means of a continuous variable are conditional on the universal set of all numbers (i.e., before the nature of a study is known and a…

Methodology · Statistics 2025-06-05 Huw Llewelyn

We study Bayesian estimation of finite mixture models in a general setup where the number of components is unknown and allowed to grow with the sample size. An assumption on growing number of components is a natural one as the degree of…

Statistics Theory · Mathematics 2022-03-18 Ilsang Ohn , Lizhen Lin

When scholars suspect units are dependent on each other within clusters but independent of each other across clusters, they employ cluster-robust standard errors (CRSEs). Nevertheless, what to cluster over is sometimes unknown. For…

Methodology · Statistics 2025-11-12 Kentaro Fukumoto

Approximate Bayesian computation (ABC) has become an essential part of the Bayesian toolbox for addressing problems in which the likelihood is prohibitively expensive or entirely unknown, making it intractable. ABC defines a…

Methodology · Statistics 2020-07-14 Hien D. Nguyen , Julyan Arbel , Hongliang Lü , Florence Forbes

Optimal implementation and monitoring of wind energy generation hinge on reliable power modeling that is vital for understanding turbine control, farm operational optimization, and grid load balance. Based on the idea of similar wind…

Machine Learning · Computer Science 2022-04-05 Hao Chen

Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a…

Methodology · Statistics 2025-10-21 Fumiya Iwashige , Tomoya Wakayama , Shonosuke Sugasawa , Shintaro Hashimoto

Choosing appropriate hyperparameters for unsupervised clustering algorithms in an optimal way depending on the problem under study is a long standing challenge, which we tackle while adapting clustering algorithms for immune disorder…

Quantitative Methods · Quantitative Biology 2020-09-25 A. Carpio , A. Simón , L. F. Villa

Bayesian clustering typically relies on mixture models, with each component interpreted as a different cluster. After defining a prior for the component parameters and weights, Markov chain Monte Carlo (MCMC) algorithms are commonly used to…

Methodology · Statistics 2024-07-30 Alexander Dombowsky , David B. Dunson

Star clusters are studied widely both as benchmarks for stellar evolution models and in their own right. Cluster age and mass distributions within galaxies are probes of star formation histories, and of cluster formation and disruption…

Solar and Stellar Astrophysics · Physics 2015-05-18 Morgan Fouesneau , Ariane Lançon

We address the problem of data clustering by introducing an unsupervised, parameter free approach based on maximum likelihood principle. Starting from the observation that data sets belonging to the same cluster share a common information,…

Statistical Mechanics · Physics 2009-11-07 Lorenzo Giada , Matteo Marsili

Unsupervised clustering, also known as natural clustering, stands for the classification of data according to their similarities. Here we study this problem from the perspective of complex networks. Mapping the description of data…

Data Analysis, Statistics and Probability · Physics 2012-08-22 Clara Granell , Sergio Gomez , Alex Arenas

In recent years, large-scale Bayesian learning draws a great deal of attention. However, in big-data era, the amount of data we face is growing much faster than our ability to deal with it. Fortunately, it is observed that large-scale…

Machine Learning · Computer Science 2022-02-15 Qianqian Song

The use of a finite mixture of normal distributions in model-based clustering allows to capture non-Gaussian data clusters. However, identifying the clusters from the normal components is challenging and in general either achieved by…

Methodology · Statistics 2016-06-21 Gertraud Malsiner-Walli , Sylvia Frühwirth-Schnatter , Bettina Grün