English
Related papers

Related papers: Reducing over-clustering via the powered Chinese r…

200 papers

Finite mixture models are flexible methods that are commonly used for model-based clustering. A recent focus in the model-based clustering literature is to highlight the difference between the number of components in a mixture model and the…

Methodology · Statistics 2023-08-03 Garritt L. Page , Massimo Ventrucci , Maria Franco-Villoria

Clustering mixed data presents numerous challenges inherent to the very heterogeneous nature of the variables. A clustering algorithm should be able, despite of this heterogeneity, to extract discriminant pieces of information from the…

Machine Learning · Computer Science 2022-05-10 Robin Fuchs , Denys Pommeret , Cinzia Viroli

We propose a restricted collapsed draw (RCD) sampler, a general Markov chain Monte Carlo sampler of simultaneous draws from a hierarchical Chinese restaurant process (HCRP) with restriction. Models that require simultaneous draws from a…

Machine Learning · Statistics 2011-06-03 Takaki Makino , Shunsuke Takei , Issei Sato , Daichi Mochihashi

In model-based-clustering mixture models are used to group data points into clusters. A useful concept introduced for Gaussian mixtures by Malsiner Walli et al (2016) are sparse finite mixtures, where the prior distribution on the weight…

Methodology · Statistics 2018-08-23 Sylvia Frühwirth-Schnatter , Gertraud Malsiner-Walli

The Pitman-Yor, or Chinese Restaurant Process, is a stochastic process that generates distributions following a power-law with exponents lower than two, as found in a numerous physical, biological, technological and social systems. We…

Disordered Systems and Neural Networks · Physics 2015-05-13 Bruno Bassetti , Mina Zarei , Marco Cosentino Lagomarsino , Ginestra Bianconi

Short text clustering has become increasingly important with the popularity of social media like Twitter, Google+, and Facebook. Existing methods can be broadly categorized into two paradigms: topic model-based approaches and deep…

Computation and Language · Computer Science 2025-07-21 Enhao Cheng , Shoujia Zhang , Jianhua Yin , Xuemeng Song , Tian Gan , Liqiang Nie

Network data often represent multiple types of relations, which can also denote exchanged quantities, and are typically encompassed in a weighted multiplex. Such data frequently exhibit clustering structures, however, traditional clustering…

Methodology · Statistics 2024-12-17 Iuliia Promskaia , Adrian O'Hagan , Michael Fop

We present a novel approach, in which we learn to cluster data directly from side information, in the form of a small set of pairwise examples. Unlike previous methods, with or without side information, we do not need to know the number of…

Machine Learning · Computer Science 2023-05-31 Michael A. Hobley , Victor A. Prisacariu

DP-means clustering was obtained as an extension of $K$-means clustering. While it is implemented with a simple and efficient algorithm, it can estimate the number of clusters simultaneously. However, DP-means is specifically designed for…

Machine Learning · Computer Science 2021-08-26 Masahiro Kobayashi , Kazuho Watanabe

Employing nonparametric methods for density estimation has become routine in Bayesian statistical practice. Models based on discrete nonparametric priors such as Dirichlet Process Mixture (DPM) models are very attractive choices due to…

Methodology · Statistics 2017-07-03 J. J. Quinlan , F. A. Quintana , G. L. Page

We establish scaling limit theorems for the up-down ordered Chinese restaurant processes (oCRPs) of Rogers and Winkel as processes in a space of interval partitions. As previously conjectured, the limits are self-similar diffusions…

Probability · Mathematics 2025-12-09 Quan Shi , Matthias Winkel

Many popular random partition models, such as the Chinese restaurant process and its two-parameter extension, fall in the class of exchangeable random partitions, and have found wide applicability in model-based clustering, population…

Methodology · Statistics 2017-11-21 Giuseppe Di Benedetto , François Caron , Yee Whye Teh

Clustering multi-dimensional points is a fundamental task in many fields, and density-based clustering supports many applications as it can discover clusters of arbitrary shapes. This paper addresses the problem of Density-Peaks Clustering…

Databases · Computer Science 2022-12-01 Daichi Amagata , Takahiro Hara

Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process…

Mixtures of multivariate normal inverse Gaussian (MNIG) distributions can be used to cluster data that exhibit features such as skewness and heavy tails. However, for cluster analysis, using a traditional finite mixture model framework,…

Methodology · Statistics 2020-05-13 Yuan Fang , Dimitris Karlis , Sanjeena Subedi

The hospitality industry is one of the data-rich industries that receives huge Volumes of data streaming at high Velocity with considerably Variety, Veracity, and Variability. These properties make the data analysis in the hospitality…

Databases · Computer Science 2017-09-20 Avishek Bose , Arslan Munir , Neda Shabani

Recently, diffusion probabilistic models (DPMs) have achieved promising results in diverse generative tasks. A typical DPM framework includes a forward process that gradually diffuses the data distribution and a reverse process that…

Machine Learning · Computer Science 2023-10-31 Tianyu Pang , Cheng Lu , Chao Du , Min Lin , Shuicheng Yan , Zhijie Deng

To identify novel dynamic patterns of gene expression, we develop a statistical method to cluster noisy measurements of gene expression collected from multiple replicates at multiple time points, with an unknown number of clusters. We…

Applications · Statistics 2013-12-02 Audrey Qiuyan Fu , Steven Russell , Sarah J. Bray , Simon Tavaré

Cluster analysis plays an important role in decision making process for many knowledge-based systems. There exist a wide variety of different approaches for clustering applications including the heuristic techniques, probabilistic models,…

Artificial Intelligence · Computer Science 2017-03-09 Kayvan Bijari , Hadi Zare , Hadi Veisi , Hossein Bobarshad

We investigate a disordered variant of Pitman's Chinese restaurant process where tables carry i.i.d. weights. Incoming customers choose to sit at an occupied table with a probability proportional to the product of its occupancy and its…

Probability · Mathematics 2024-05-06 Jakob E. Björnberg , Cécile Mailler , Peter Mörters , Daniel Ueltschi