中文
相关论文

相关论文: Cluster and Feature Modeling from Combinatorial St…

200 篇论文

The use of a finite mixture of normal distributions in model-based clustering allows to capture non-Gaussian data clusters. However, identifying the clusters from the normal components is challenging and in general either achieved by…

统计方法学 · 统计学 2016-06-21 Gertraud Malsiner-Walli , Sylvia Frühwirth-Schnatter , Bettina Grün

The natural habitat of most Bayesian methods is data represented by exchangeable sequences of observations, for which de Finetti's theorem provides the theoretical foundation. Dirichlet process clustering, Gaussian process regression, and…

统计理论 · 数学 2015-02-16 Peter Orbanz , Daniel M. Roy

Traditional Bayesian random partition models assume that the size of each cluster grows linearly with the number of data points. While this is appealing for some applications, this assumption is not appropriate for other tasks such as…

统计方法学 · 统计学 2020-04-07 Brenda Betancourt , Giacomo Zanella , Rebecca C. Steorts

A model-based approach is developed for clustering categorical data with no natural ordering. The proposed method exploits the Hamming distance to define a family of probability mass functions to model the data. The elements of this family…

统计方法学 · 统计学 2024-07-02 Raffaele Argiento , Edoardo Filippi-Mazzola , Lucia Paci

In this paper we propose a Bayesian nonparametric model for clustering partial ranking data. We start by developing a Bayesian nonparametric extension of the popular Plackett-Luce choice model that can handle an infinite number of choice…

机器学习 · 统计学 2014-08-04 François Caron , Yee Whye Teh , Thomas Brendan Murphy

Recent work on overfitting Bayesian mixtures of distributions offers a powerful framework for clustering multivariate data using a latent Gaussian model which resembles the factor analysis model. The flexibility provided by overfitting…

统计方法学 · 统计学 2019-08-29 Panagiotis Papastamoulis

In this chapter we review some examples, methods, and recent results involving comparison of clustering properties of point processes. Our approach is founded on some basic observations allowing us to consider void probabilities and moment…

概率论 · 数学 2014-05-23 Bartłomiej Błaszczyszyn , D. Yogeshwaran

We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a…

Bayesian nonparametrics are a class of probabilistic models in which the model size is inferred from data. A recently developed methodology in this field is small-variance asymptotic analysis, a mathematical technique for deriving learning…

机器学习 · 统计学 2017-07-27 Trevor Campbell , Brian Kulis , Jonathan How

The parsimonious Gaussian mixture models, which exploit an eigenvalue decomposition of the group covariance matrices of the Gaussian mixture, have shown their success in particular in cluster analysis. Their estimation is in general…

机器学习 · 统计学 2018-10-18 Faicel Chamroukhi , Marius Bartcus , Hervé Glotin

We consider the Bayesian mixture of finite mixtures (MFMs) and Dirichlet process mixture (DPM) models for clustering. Recent asymptotic theory has established that DPMs overestimate the number of clusters for large samples and that…

统计方法学 · 统计学 2022-08-01 Yannis Chaumeny , Johan van der Molen Moris , Anthony C. Davison , Paul D. W. Kirk

Functional concurrent, or varying-coefficient, regression models are commonly used in biomedical and clinical settings to investigate how the relation between an outcome and observed covariate varies as a function of another covariate. In…

统计方法学 · 统计学 2024-10-10 Mingrui Liang , Matthew D. Koslovsky , Emily T. Hebert , Darla E. Kendzor , Marina Vannucci

Clustering is one of the most widely used procedures in the analysis of microarray data, for example with the goal of discovering cancer subtypes based on observed heterogeneity of genetic marks between different tissues. It is well-known…

统计方法学 · 统计学 2009-04-21 Heng Lian

Joint alignment of a collection of functions is the process of independently transforming the functions so that they appear more similar to each other. Typically, such unsupervised alignment algorithms fail when presented with complex data…

机器学习 · 计算机科学 2012-10-19 Marwan A. Mattar , Allen R. Hanson , Erik G. Learned-Miller

Neyman-Scott processes (NSPs) are point process models that generate clusters of points in time or space. They are natural models for a wide range of phenomena, ranging from neural spike trains to document streams. The clustering property…

机器学习 · 统计学 2023-09-13 Yixin Wang , Anthony Degleris , Alex H. Williams , Scott W. Linderman

The problem of inferring a clustering of a data set has been the subject of much research in Bayesian analysis, and there currently exists a solid mathematical foundation for Bayesian approaches to clustering. In particular, the class of…

概率论 · 数学 2013-01-30 Tamara Broderick , Jim Pitman , Michael I. Jordan

Clustering has become a core technology in machine learning, largely due to its application in the field of unsupervised learning, clustering, classification, and density estimation. A frequentist approach exists to hand clustering based on…

机器学习 · 计算机科学 2021-08-27 Jun Lu

We present a consensus Monte Carlo algorithm that scales existing Bayesian nonparametric models for clustering and feature allocation to big data. The algorithm is valid for any prior on random subsets such as partitions and latent feature…

统计计算 · 统计学 2020-02-26 Yang Ni , Yuan Ji , Peter Mueller

We conduct cluster analysis on a class of locally asymptotically self-similar stochastic processes, which includes multifractional Brownian motion as a representative. When the true number of clusters is supposed to be known, a new…

机器学习 · 统计学 2020-01-15 Qidi Peng , Nan Rao , Ran Zhao

The problem of time-series clustering is considered in the case where each data-point is a sample generated by a piecewise stationary ergodic process. Stationary processes are perhaps the most general class of processes considered in…

机器学习 · 统计学 2019-06-27 Azadeh Khaleghi , Daniil Ryabko