中文
相关论文

相关论文: Powered Dirichlet Process for Controlling the Impo…

200 篇论文

Prior distributions play a crucial role in Bayesian approaches to clustering. Two commonly-used prior distributions are the Dirichlet and Pitman-Yor processes. In this paper, we investigate the predictive probabilities that underlie these…

统计方法学 · 统计学 2010-10-18 Hanna M. Wallach , Shane T. Jensen , Lee Dicker , Katherine A. Heller

Dirichlet process mixture (DPM) models tend to produce many small clusters regardless of whether they are needed to accurately characterize the data - this is particularly true for large data sets. However, interpretability, parsimony, data…

机器学习 · 计算机科学 2018-02-16 Jun Lu , Meng Li , David Dunson

The Bayesian approach to inference stands out for naturally allowing borrowing information across heterogeneous populations, with different samples possibly sharing the same distribution. A popular Bayesian nonparametric model for…

统计方法学 · 统计学 2022-01-25 Antonio Lijoi , Igor Prünster , Giovanni Rebaudo

The two parameter Poisson-Dirichlet Process (PDP), a generalisation of the Dirichlet Process, is increasingly being used for probabilistic modelling in discrete areas such as language technology, bioinformatics, and image analysis. There is…

统计理论 · 数学 2012-02-17 Wray Buntine , Marcus Hutter

We develop a Bayesian framework for tackling the supervised clustering problem, the generic problem encountered in tasks such as reference matching, coreference resolution, identity uncertainty and record linkage. Our clustering model is…

机器学习 · 计算机科学 2009-07-07 Hal Daumé , Daniel Marcu

For a long time, the Dirichlet process has been the gold standard discrete random measure in Bayesian nonparametrics. The Pitman--Yor process provides a simple and mathematically tractable generalization, allowing for a very flexible…

统计理论 · 数学 2020-01-08 Caroline Lawless , Julyan Arbel

Dirichlet process mixtures are particularly sensitive to the value of the precision parameter controlling the behavior of the latent partition. Randomization of the precision through a prior distribution is a common solution, which leads to…

统计方法学 · 统计学 2024-09-04 Alessandro Zito , Tommaso Rigon , David B. Dunson

The hierarchical Dirichlet process is the cornerstone of Bayesian nonparametric multilevel models. Its generative model can be described through a set of latent variables, commonly referred to as tables within the popular restaurant…

统计理论 · 数学 2025-05-06 Marta Catalano , Claudio Del Sole

We present the nested Chinese restaurant process (nCRP), a stochastic process which assigns probability distributions to infinitely-deep, infinitely-branching trees. We show how this stochastic process can be used as a prior distribution in…

机器学习 · 统计学 2009-08-27 David M. Blei , Thomas L. Griffiths , Michael I. Jordan

Finite mixture models are flexible methods that are commonly used for model-based clustering. A recent focus in the model-based clustering literature is to highlight the difference between the number of components in a mixture model and the…

统计方法学 · 统计学 2023-08-03 Garritt L. Page , Massimo Ventrucci , Maria Franco-Villoria

Factorized Information Criterion (FIC) is a recently developed information criterion, based on which a novel model selection methodology, namely Factorized Asymptotic Bayesian (FAB) Inference, has been developed and successfully applied to…

机器学习 · 统计学 2015-07-02 Shaohua Li

The distance dependent Chinese Restaurant Process (ddCRP) provides a flexible prior distribution for clustering observations, incorporating covariate information through pairwise distances and accommodating a rich variety of cluster…

统计方法学 · 统计学 2026-05-18 Joseph Marsh , Theodore Kypraios , Rowland G. Seymour

The advent of Generative Artificial Intelligence (GAI) has heralded an inflection point that changed how society thinks about knowledge acquisition. While GAI cannot be fully trusted for decision-making, it may still provide valuable…

统计方法学 · 统计学 2025-05-20 Sean O'Hagan , Veronika Ročková

This article proposes a Bayesian nonparametric method for forecasting, imputation, and clustering in sparsely observed, multivariate time series data. The method is appropriate for jointly modeling hundreds of time series with widely…

统计方法学 · 统计学 2019-02-27 Feras A. Saad , Vikash K. Mansinghka

One of the focal points of the modern literature on Bayesian nonparametrics has been the problem of clustering, or partitioning, where each data point is modeled as being associated with one and only one of some collection of groups called…

统计理论 · 数学 2013-10-02 Tamara Broderick , Michael I. Jordan , Jim Pitman

Exemplar-based clustering methods have been shown to produce state-of-the-art results on a number of synthetic and real-world clustering problems. They are appealing because they offer computational benefits over latent-mean models and can…

机器学习 · 计算机科学 2012-06-18 Daniel Tarlow , Richard S. Zemel , Brendan J. Frey

While there is an immense literature on Bayesian methods for clustering, the multiview case has received little attention. This problem focuses on obtaining distinct but statistically dependent clusterings in a common set of entities for…

统计方法学 · 统计学 2025-04-01 Alexander Dombowsky , David B. Dunson

Dirichlet distribution and Dirichlet process as its infinite dimensional generalization are primarily used conjugate prior of categorical and multinomial distributions in Bayesian statistics. Extensions have been proposed to broaden…

统计方法学 · 统计学 2014-12-05 Xuenan Feng

Tree structures are ubiquitous in data across many domains, and many datasets are naturally modelled by unobserved tree structures. In this paper, first we review the theory of random fragmentation processes [Bertoin, 2006], and a number of…

机器学习 · 统计学 2015-09-17 Hong Ge , Yarin Gal , Zoubin Ghahramani

We consider the problem of clustering grouped data with possibly non-exchangeable groups whose dependencies can be characterized by a known directed acyclic graph. To allow the sharing of clusters among the non-exchangeable groups, we…

‹ 上一页 1 2 3 10 下一页 ›