English
Related papers

Related papers: Replica analysis of Bayesian data clustering

200 papers

In order to handle large data sets omnipresent in modern science, efficient compression algorithms are necessary. Here, a Bayesian data compression (BDC) algorithm that adapts to the specific measurement situation is derived in the context…

Data Analysis, Statistics and Probability · Physics 2021-03-01 Johannes Harth-Kitzerow , Reimar Leike , Philipp Arras , Torsten A. Enßlin

In this paper, we consider the problem of partitioning a small data sample of size $n$ drawn from a mixture of $2$ sub-gaussian distributions. Our work is motivated by the application of clustering individuals according to their population…

Statistics Theory · Mathematics 2023-01-05 Shuheng Zhou

The paper tackles the problem of clustering multiple networks, directed or not, that do not share the same set of vertices, into groups of networks with similar topology. A statistical model-based approach based on a finite mixture of…

Statistics Theory · Mathematics 2023-11-07 Tabea Rebafka

Using computer simulations, we validate a simple free energy model that can be analytically solved to predict the equilibrium size of self-limiting clusters of particles in the fluid state governed by a combination of short-range attractive…

Soft Condensed Matter · Physics 2016-08-25 Jonathan A. Bollinger , Thomas M. Truskett

We discuss the development of cluster algorithms from the viewpoint of probability theory and not from the usual viewpoint of a particular model. By using the perspective of probability theory, we detail the nature of a cluster algorithm,…

Condensed Matter · Physics 2009-10-22 N. Kawashima , J. E. Gubernatis

Increased deployment of residential smart meters has made it possible to record energy consumption data on short intervals. These data, if used efficiently, carry valuable information for managing power demand and increasing energy…

Other Computer Science · Computer Science 2019-03-05 Nameer Al Khafaf , Mahdi Jalili , Peter Sokolowski

We build networks of genetic similarity in which the nodes are organisms sampled from biological populations. The procedure is illustrated by constructing networks from genetic data of a marine clonal plant. An important feature in the…

Populations and Evolution · Quantitative Biology 2008-01-23 E. Hernandez-Garcia , A. F. Rozenfeld , V. M. Eguiluz , S. Arnaud-Haond , C. M. Duarte

Bayesian nonparametrics are a class of probabilistic models in which the model size is inferred from data. A recently developed methodology in this field is small-variance asymptotic analysis, a mathematical technique for deriving learning…

Machine Learning · Statistics 2017-07-27 Trevor Campbell , Brian Kulis , Jonathan How

Robustly determining the optimal number of clusters in a data set is an essential factor in a wide range of applications. Cluster enumeration becomes challenging when the true underlying structure in the observed data is corrupted by…

Signal Processing · Electrical Eng. & Systems 2021-05-06 Christian A. Schroth , Michael Muma

We propose a MAP Bayesian approach to perform and evaluate a co-clustering of mixed-type data tables. The proposed model infers an optimal segmentation of all variables then performs a co-clustering by minimizing a Bayesian model selection…

Machine Learning · Statistics 2019-02-07 Aichetou Bouchareb , Marc Boullé , Fabrice Rossi , Fabrice Clérot

We study the free energy of a most used deep architecture for restricted Boltzmann machines, where the layers are disposed in series. Assuming independent Gaussian distributed random weights, we show that the error term in the so-called…

Disordered Systems and Neural Networks · Physics 2022-07-08 Giuseppe Genovese

Recent advances in engineering technologies have enabled the collection of a large number of longitudinal features. This wealth of information presents unique opportunities for researchers to investigate the complex nature of diseases and…

Methodology · Statistics 2023-11-27 Zihang Lu , Noirrit Kiran Chandra

We study statistical and computational limits of clustering when the means of the centres are sparse and their dimension is possibly much larger than the sample size. Our theoretical analysis focuses on the model $X_i = z_i \theta +…

Statistics Theory · Mathematics 2021-03-23 Matthias Löffler , Alexander S. Wein , Afonso S. Bandeira

We consider an extension of model-based clustering to the semi-supervised case, where some of the data are pre-labeled. We provide a derivation of the Bayesian Information Criterion (BIC) approximation to the Bayes factor in this setting.…

Methodology · Statistics 2016-04-28 Jordan Yoder , Carey E. Priebe

Consider a population of organisms that harvest free energy from their environment to reproduce. This paper shows that if the organisms' reproductive rates are proportional to the amount of physical free energy that they can convert into…

Populations and Evolution · Quantitative Biology 2025-11-25 Seth Lloyd

Our goal is to deploy a high-accuracy system starting with zero training examples. We consider an "on-the-job" setting, where as inputs arrive, we use real-time crowdsourcing to resolve uncertainty where needed and output our prediction…

Artificial Intelligence · Computer Science 2015-12-09 Keenon Werling , Arun Chaganty , Percy Liang , Chris Manning

Describing the complex dependence structure of extreme phenomena is particularly challenging. To tackle this issue we develop a novel statistical algorithm that describes extremal dependence taking advantage of the inherent hierarchical…

Methodology · Statistics 2018-07-24 Sabrina Vettori , Raphaël Huser , Johan Segers , Marc G. Genton

Partially recorded data are frequently encountered in many applications and usually clustered by first removing incomplete cases or features with missing values, or by imputing missing values, followed by application of a clustering…

Methodology · Statistics 2021-10-20 Emily M. Goren , Ranjan Maitra

A computational theory for clustering and a semi-supervised clustering algorithm is presented. Clustering is defined to be the obtainment of groupings of data such that each group contains no anomalies with respect to a chosen grouping…

Machine Learning · Computer Science 2025-07-17 Nassir Mohammad

We propose a score function for Bayesian clustering. The function is parameter free and captures the interplay between the within cluster variance and the between cluster entropy of a clustering. It can be used to choose the number of…

Other Statistics · Statistics 2019-05-27 John Noble , Łukasz Rajkowski