English
Related papers

Related papers: Dirichlet-tree multinomial mixtures for clustering…

200 papers

We develop a dependent Dirichlet process (DDP) model for repeated measures multiple membership (MM) data. This data structure arises in studies under which an intervention is delivered to each client through a sequence of elements which…

Applications · Statistics 2013-12-09 Terrance D. Savitsky , Susan M. Paddock

There is a rich literature on Bayesian methods for density estimation, which characterize the unknown density as a mixture of kernels. Such methods have advantages in terms of providing uncertainty quantification in estimation, while being…

Methodology · Statistics 2024-04-10 Shounak Chattopadhyay , Antik Chakraborty , David B. Dunson

Recent advances in bioinformatics have made high-throughput microbiome data widely available, and new statistical tools are required to maximize the information gained from these data. For example, analysis of high-dimensional microbiome…

Methodology · Statistics 2017-03-23 Neal S. Grantham , Brian J. Reich , Elizabeth T. Borer , Kevin Gross

It is crucial to provide compatible treatment schemes for a disease according to various symptoms at different stages. However, most classification methods might be ineffective in accurately classifying a disease that holds the…

Machine Learning · Computer Science 2019-11-26 Jianguo Chen , Kenli Li , Huigui Rong , Kashif Bilal , Nan Yang , Keqin Li

Do expert-defined or diagnostically-labeled data groups align with clusters inferred through statistical modeling? If not, where do discrepancies between predefined labels and model-based groupings occur and why? In this work, we introduce…

Methodology · Statistics 2026-03-18 Patricia Puchhammer , Ines Wilms , Peter Filzmoser

Clustering algorithms start with a fixed divergence, which captures the possibly asymmetric distance between a sample and a centroid. In the mixture model setting, the sample distribution plays the same role. When all attributes have the…

Machine Learning · Computer Science 2017-01-10 Mehmet Emin Basbug , Barbara Engelhardt

Clustering is essential in data analysis and machine learning, but traditional algorithms like $k$-means and Gaussian Mixture Models (GMM) often fail with nonconvex clusters. To address the challenge, we introduce the Flexible Bivariate…

Machine Learning · Computer Science 2025-02-28 Yung-Peng Hsu , Hung-Hsuan Chen

Clustering has become a core technology in machine learning, largely due to its application in the field of unsupervised learning, clustering, classification, and density estimation. A frequentist approach exists to hand clustering based on…

Machine Learning · Computer Science 2021-08-27 Jun Lu

Phylogenetic inference can potentially result in a more accurate tree using data from multiple loci. However, if the loci are incongruent--due to events such as incomplete lineage sorting or horizontal gene transfer--it can be misleading to…

Populations and Evolution · Quantitative Biology 2016-03-10 Kevin Gori , Tomasz Suchan , Nadir Alvarez , Nick Goldman , Christophe Dessimoz

Relative abundance is a common metric to estimate the composition of species in ecological surveys reflecting patterns of commonness and rarity of biological assemblages. Measurements of coral reef compositions formed by four communities…

Applications · Statistics 2021-05-06 Luiza Piancastelli , Nial Friel , Julie Vercelloni , Kerrie Mengersen , Antonietta Mira

In Bayesian inference for mixture models with an unknown number of components, a finite mixture model is usually employed that assumes prior distributions for mixing weights and the number of components. This model is called a mixture of…

Methodology · Statistics 2025-12-25 Fumiya Iwashige , Shintaro Hashimoto

Clustering is commonly performed as an initial analysis step for uncovering structure in 'omics datasets, e.g. to discover molecular subtypes of disease. The high-throughput, high-dimensional nature of these datasets means that they provide…

Methodology · Statistics 2023-03-02 Paul D. W. Kirk , Filippo Pagani , Sylvia Richardson

A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity…

Machine Learning · Computer Science 2019-06-18 Amir Najafi , Abolfazl Motahari , Hamid R. Rabiee

Mixture models are commonly used in applications with heterogeneity and overdispersion in the population, as they allow the identification of subpopulations. In the Bayesian framework, this entails the specification of suitable prior…

Methodology · Statistics 2023-06-21 Andrea Cremaschi , Timothy M. Wertz , Maria De Iorio

Dirichlet Process Mixture (DPM) models have been increasingly employed to specify random partition models that take into account possible patterns within the covariates. Furthermore, to deal with large numbers of covariates, methods for…

Applications · Statistics 2016-11-01 William Barcella , Maria De Iorio , Gianluca Baio

Recent evidence suggests that analyzing the presence/absence of taxonomic features can offer a compelling alternative to differential abundance analysis in microbiome studies. However, standard approaches to differential prevalence analysis…

Methodology · Statistics 2026-05-26 Juho Pelto , Kari Auranen , Janne V. Kujala , Leo Lahti

We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We…

Machine Learning · Statistics 2010-07-16 Lauren A. Hannah , David M. Blei , Warren B. Powell

In this article, we develop a new class of multivariate distributions adapted for count data, called Tree P\'olya Splitting. This class results from the combination of a univariate distribution and singular multivariate distributions along…

Statistics Theory · Mathematics 2025-01-30 Samuel Valiquette , Jean Peyhardi , Éric Marchand , Gwladys Toulemonde , Frédéric Mortier

Clustering is a widely used technique in data mining applications for discovering patterns in underlying data. Most traditional clustering algorithms are limited to handling datasets that contain either numeric or categorical attributes.…

Artificial Intelligence · Computer Science 2007-05-23 Zengyou He , Xiaofei Xu , Shengchun Deng

In agglomerative hierarchical clustering, pair-group methods suffer from a problem of non-uniqueness when two or more distances between different clusters coincide during the amalgamation process. The traditional approach for solving this…

Information Retrieval · Computer Science 2009-06-10 Alberto Fernandez , Sergio Gomez