English
Related papers

Related papers: A Bernoulli Mixture Model to Understand and Predic…

200 papers

The progression of chronic diseases often follows highly variable trajectories, and the underlying factors remain poorly understood. Standard mixed-effects models typically represent inter-patient differences as random deviations around a…

We study the large sample behavior of a convex clustering framework, which minimizes the sample within cluster sum of squares under an~$\ell_1$ fusion constraint on the cluster centroids. This recently proposed approach has been gaining in…

Methodology · Statistics 2016-12-30 Peter Radchenko , Gourab Mukherjee

Childhood anemia remains a major public health challenge in Nepal and is associated with impaired growth, cognition, and increased morbidity. Using World Health Organization hemoglobin thresholds, we defined anemia status for children aged…

Machine Learning · Computer Science 2026-02-03 Deepak Bastola , Pitambar Acharya , Dipak Dulal , Rabina Dhakal , Yang Li

Clustering multivariate binary data is of interest in many scientific fields, including ecology, biomedicine, and social policy. Beyond heuristic clustering algorithms, such data can be modelled using multivariate Bernoulli mixture models.…

Methodology · Statistics 2026-04-24 Luisa Ferrari , Maria Franco Villoria , Garritt L. Page , Alex Laini

Motivated by the problem of accurately predicting gap times between successive blood donations, we present here a general class of Bayesian nonparametric models for clustering. These models allow for prediction of new recurrences,…

Methodology · Statistics 2022-10-18 Raffaele Argiento , Riccardo Corradin , Alessandra Guglielmi , Ettore Lanzarone

Clustering is a pivotal challenge in unsupervised machine learning and is often investigated through the lens of mixture models. The optimal error rate for recovering cluster labels in Gaussian and sub-Gaussian mixture models involves ad…

Statistics Theory · Mathematics 2024-07-18 Maximilien Dreveton , Alperen Gözeten , Matthias Grossglauser , Patrick Thiran

Objectives: Unsupervised learning with electronic health record (EHR) data has shown promise for phenotype discovery, but approaches typically disregard existing clinical information, limiting interpretability. We operationalize a Bayesian…

Background:Adverse reproductive history is a multisystemic risk factor, but evidence is constrained by isolated outcome studies, limited adjustment, and non-interpretable algorithmic models. We re-frame the estimand from prediction to…

Other Quantitative Biology · Quantitative Biology 2026-04-28 Sunday A. Adetunji

In microbiome studies, it is often of great interest to identify clusters or partitions of microbiome profiles within a study population and to characterize the distinctive attributes of each resulting microbial community. While raw counts…

Methodology · Statistics 2025-08-18 Zhongmao Liu , Xiaohui Yin , Yanjiao Zhou , Gen Li , Kun Chen

Model-based clustering methods for continuous data are well established and commonly used in a wide range of applications. However, model-based clustering methods for categorical data are less standard. Latent class analysis is a commonly…

Methodology · Statistics 2013-02-20 Isabella Gollini , Thomas Brendan Murphy

In this paper, we first study the fundamental limit of clustering networks when a multi-layer network is present. Under the mixture multi-layer stochastic block model (MMSBM), we show that the minimax optimal network clustering error rate,…

Statistics Theory · Mathematics 2023-11-28 Zhongyuan Lyu , Ting Li , Dong Xia

Undernutrition, resulting in restricted growth, and quantified here using height-for-age z-scores, is an important contributor to childhood morbidity and mortality. Since all levels of mild, moderate and severe undernutrition are of…

Applications · Statistics 2015-12-08 Mariel M. Finucane , Christopher J. Paciorek , Gretchen A. Stevens , Majid Ezzati

Unsupervised learning is often used to uncover clusters in data. However, different kinds of noise may impede the discovery of useful patterns from real-world time-series data. In this work, we focus on mitigating the interference of…

Machine Learning · Statistics 2021-12-07 Irene Y. Chen , Rahul G. Krishnan , David Sontag

Data represented as covariance-type matrices arise in many fields, including brain functional connectivity and diffusion tensor imaging. We develop the MFM-Wishart, a Bayesian model-based clustering approach for such data that combines…

Methodology · Statistics 2026-05-25 Zongyu Li , Stefano Castruccio , Zhiyong Zhang

Epilepsy affects 50 million people worldwide and is one of the most common serious neurological disorders. Seizure detection and classification is a valuable tool for diagnosing and maintaining the condition. An automated classification…

Signal Processing · Electrical Eng. & Systems 2022-08-22 Scot Davidson , Niamh McCallan , Kok Yew Ng , Pardis Biglarbeigi , Dewar Finlay , Boon Leong Lan , James McLaughlin

We study clustering methods for binary data, first defining aggregation criteria that measure the compactness of clusters. Five new and original methods are introduced, using neighborhoods and population behavior combinatorial optimization…

An estimated 300 million people worldwide suffer from asthma, and this number is expected to increase to 400 million by 2025. Approximately 250,000 people die prematurely each year from asthma out of which, almost all deaths are avoidable.…

Computers and Society · Computer Science 2018-04-13 Saksham Kukreja

The aim of this study is to show the importance of two classification techniques, viz. decision tree and clustering, in prediction of learning disabilities (LD) of school-age children. LDs affect about 10 percent of all children enrolled in…

Artificial Intelligence · Computer Science 2010-11-03 Julie M. David And Kannan Balakrishnan

Clustered data is ubiquitous in a variety of scientific fields. In this paper, we propose a flexible and interpretable modeling approach, called grouped heterogenous mixture modeling, for clustered data, which models cluster-wise…

Methodology · Statistics 2020-02-10 Shonosuke Sugasawa

We introduce a new approach to deciding the number of clusters. The approach is applied to Optimally Tuned Robust Improper Maximum Likelihood Estimation (OTRIMLE; Coretto and Hennig 2016) of a Gaussian mixture model allowing for…

Methodology · Statistics 2020-12-29 Christian Hennig , Pietro Coretto