中文
相关论文

相关论文: A Bayesian latent allocation model for clustering …

200 篇论文

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

统计方法学 · 统计学 2025-05-16 Luca Scrucca

We develop a Bayesian framework for tackling the supervised clustering problem, the generic problem encountered in tasks such as reference matching, coreference resolution, identity uncertainty and record linkage. Our clustering model is…

机器学习 · 计算机科学 2009-07-07 Hal Daumé , Daniel Marcu

Biogeographical regions (geographically distinct assemblages of species and communities) constitute a cornerstone for ecology, biogeography, evolution and conservation biology. Species turnover measures are often used to quantify…

定量方法 · 定量生物学 2015-08-19 Daril A. Vilhena , Alexandre Antonelli

The estimation of voting blocs is an important statistical inquiry in political science. However, the scope of these analyses is usually restricted to roll call data where individual votes are directly observed. Here, we examine a Bayesian…

应用统计 · 统计学 2023-01-09 John D. O'Brien

Latent class analysis is used to perform model based clustering for multivariate categorical responses. Selection of the variables most relevant for clustering is an important task which can affect the quality of clustering considerably.…

统计计算 · 统计学 2016-06-17 Arthur White , Jason Wyse , Thomas Brendan Murphy

In model-based clustering using finite mixture models, it is a significant challenge to determine the number of clusters (cluster size). It used to be equal to the number of mixture components (mixture size); however, this may not be valid…

机器学习 · 计算机科学 2020-07-16 Shunki Kyoya , Kenji Yamanishi

Generative approaches to clustering provide information on geometric properties of clusters, whereas discriminative approaches provide boundaries between clusters. Ideas from both approaches are incorporated to present a fully unsupervised,…

机器学习 · 统计学 2026-04-28 Mackenzie R. Neal , Paul D. McNicholas , Arthur White

Loss-based clustering methods, such as k-means and its variants, are standard tools for finding groups in data. However, the lack of quantification of uncertainty in the estimated clusters is a disadvantage. Model-based clustering based on…

统计方法学 · 统计学 2020-06-11 Tommaso Rigon , Amy H. Herring , David B. Dunson

Bayesian graphical modeling provides an appealing way to obtain uncertainty estimates when inferring network structures, and much recent progress has been made for Gaussian models. These models have been used extensively in applications to…

统计方法学 · 统计学 2012-07-06 Michael Finegold , Mathias Drton

Coral bleaching is a major concern for marine ecosystems; more than half of the world's coral reefs have either bleached or died over the past three decades. Increasing sea surface temperatures, along with various spatiotemporal…

应用统计 · 统计学 2025-11-18 Soham Sarkar , Arnab Hazra

We develop a structural framework for modeling and inferring unobserved heterogeneity in dynamic panel-data models. Unlike methods treating clustering as a descriptive device, we model heterogeneity as arising from a latent clustering…

计量经济学 · 经济学 2025-10-29 Jean-Pierre Florens , Anna Simoni

Binned data often appears in different fields of research, and it is generated after summarizing the original data in a sequence of pairs of bins (or their midpoints) and frequencies. There may exist different reasons to only provide this…

统计方法学 · 统计学 2024-09-13 Asael Fabian Martínez , Carlos Díaz-Avalos

Environmental mixture approaches do not accommodate compositional outcomes, consisting of vectors constrained onto the unit simplex. This limitation poses challenges in effectively evaluating the associations between multiple concurrent…

统计方法学 · 统计学 2025-03-31 Hachem Saddiki , Joshua L. Warren , Corina Lesseur , Elena Colicino

Model-based clustering of moderate or large dimensional data is notoriously difficult. We propose a model for simultaneous dimensionality reduction and clustering by assuming a mixture model for a set of latent scores, which are then linked…

统计方法学 · 统计学 2024-06-04 Lorenzo Ghilotti , Mario Beraha , Alessandra Guglielmi

Determining the number of clusters is a fundamental issue in data clustering. Several algorithms have been proposed, including centroid-based algorithms using the Euclidean distance and model-based algorithms using a mixture of probability…

机器学习 · 计算机科学 2024-07-30 Ryosuke Motegi , Yoichi Seki

Genes are often regulated in living cells by proteins called transcription factors (TFs) that bind directly to short segments of DNA in close proximity to specific genes. These binding sites have a conserved nucleotide appearance, which is…

统计理论 · 数学 2007-06-13 Shane T. Jensen , Jun S. Liu

We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that…

机器学习 · 计算机科学 2013-01-18 Shivakumar Vaithyanathan , Byron E Dom

Researchers and managers model ecological communities to infer the biotic and abiotic variables that shape species' ranges, habitat use, and co-occurrence which, in turn, are used to support management decisions and test ecological…

应用统计 · 统计学 2020-06-01 Trevor Hefley

In this project we are interested in performing clustering of observations such that the cluster membership is influenced by a set of predictors. To that end, we employ the Bayesian nonparameteric Common Atoms Model, which is a nested…

统计方法学 · 统计学 2025-12-11 Md Yasin Ali Parh , Jeremy T. Gaskins

In this work, the goal is to estimate the abundance of an animal population using data coming from capture-recapture surveys. We leverage the prior knowledge about the population's structure to specify a parsimonious finite mixture model…