English
Related papers

Related papers: Bayesian Clustering of Transcription Factor Bindin…

200 papers

We develop methods for efficient amortized approximate Bayesian inference over posterior distributions of probabilistic clustering models, such as Dirichlet process mixture models. The approach is based on mapping distributed,…

Machine Learning · Statistics 2018-11-27 Ari Pakman , Liam Paninski

The use of a finite mixture of normal distributions in model-based clustering allows to capture non-Gaussian data clusters. However, identifying the clusters from the normal components is challenging and in general either achieved by…

Methodology · Statistics 2016-06-21 Gertraud Malsiner-Walli , Sylvia Frühwirth-Schnatter , Bettina Grün

Background. Conventional phylogenetic clustering approaches rely on arbitrary cutpoints applied a posteriori to phylogenetic estimates. Although in practice, Bayesian and bootstrap-based clustering tend to lead to similar estimates, they…

Methodology · Statistics 2017-08-10 Luc Villandré , Aurélie Labbe , Bluma Brenner , Michel Roger , David A. Stephens

Clustering is a crucial task in various domains of knowledge, including medicine, epidemiology, genomics, environmental science, economics, and visual sciences, among others. Methodologies for inferring the number of clusters have often…

Methodology · Statistics 2025-05-26 Clara Grazian

Current models for the folding of the human genome see a hierarchy stretching down from chromosome territories, through A/B compartments and TADs (topologically-associating domains), to contact domains stabilized by cohesin and CTCF.…

Biological Physics · Physics 2020-10-02 Peter R. Cook , Davide Marenduzzo

We consider the problem of model-based clustering in the presence of many correlated, mixed continuous and discrete variables, some of which may have missing values. Discrete variables are treated with a latent continuous variable approach…

Dirichlet process mixture models (DPMM) play a central role in Bayesian nonparametrics, with applications throughout statistics and machine learning. DPMMs are generally used in clustering problems where the number of clusters is not known…

Machine Learning · Statistics 2020-10-20 Chiao-Yu Yang , Eric Xia , Nhat Ho , Michael I. Jordan

Clustering tabular data is a fundamental yet challenging problem due to heterogeneous feature types, diverse data-generating mechanisms, and the absence of transferable inductive biases across datasets. Prior-fitted networks (PFNs) have…

Machine Learning · Computer Science 2026-05-15 Tianqi Zhao , Guanyang Wang , Yan Shuo Tan , Qiong Zhang

Bayesian models based on the Dirichlet process and other stick-breaking priors have been proposed as core ingredients for clustering, topic modeling, and other unsupervised learning tasks. However, due to the flexibility of these models,…

Methodology · Statistics 2022-01-27 Ryan Giordano , Runjing Liu , Michael I. Jordan , Tamara Broderick

Biological sequences may contain patterns that are signal important biomolecular functions; a classical example is regulation of gene expression by transcription factors that bind to specific patterns in genomic promoter regions. In motif…

Applications · Statistics 2015-06-04 Luis E. Carvalho

We present ensemble methods in a machine learning (ML) framework combining predictions from five known motif/binding site exploration algorithms. For a given TF the ensemble starts with position weight matrices (PWM's) for the motif,…

Genomics · Quantitative Biology 2018-05-11 Yue Fan , Mark Kon , Charles DeLisi

The discovery of disease subtypes is an essential step for developing precision medicine, and disease subtyping via omics data has become a popular approach. While promising, subtypes obtained from conventional approaches may not be…

Applications · Statistics 2023-09-28 Lingsong Meng , Zhiguang Huo

Motivated by the need to model the dependence between regions of interest in functional neuroconnectivity for efficient inference, we propose a new sampling-based Bayesian clustering approach for covariance structures of high-dimensional…

Methodology · Statistics 2024-01-09 Hyoshin Kim , Sujit K. Ghosh , Adriana Di Martino , Emily C. Hector

Clustering is an essential technique for network analysis, with applications in a diverse range of fields. Although spectral clustering is a popular and effective method, it fails to consider higher-order structure and can perform poorly on…

Social and Information Networks · Computer Science 2020-09-14 William George Underwood , Andrew Elliott , Mihai Cucuringu

Many cellular responses to surrounding cues require temporally concerted transcriptional regulation of multiple genes. In prokaryotic cells, a single-input-module motif with one transcription factor regulating multiple target genes can…

Subcellular Processes · Quantitative Biology 2019-06-19 Jingyu Zhang , Hengyu Chen , Ruoyan Li , David A. Taft , Guang Yao , Fan Bai , Jianhua Xing

Bayesian hierarchical modeling is a natural framework to effectively integrate data and borrow information across groups. In this paper, we address problems related to density estimation and identifying clusters across related groups, by…

Methodology · Statistics 2025-10-29 Huizi Zhang , Sara Wade , Natalia Bochkina

Posterior computation in hierarchical Dirichlet process (HDP) mixture models is an active area of research in nonparametric Bayes inference of grouped data. Existing literature almost exclusively focuses on the Chinese restaurant franchise…

Computation · Statistics 2024-08-06 Snigdha Das , Yabo Niu , Yang Ni , Bani K. Mallick , Debdeep Pati

We present clustering methods for multivariate data exploiting the underlying geometry of the graphical structure between variables. As opposed to standard approaches that assume known graph structures, we first estimate the edge structure…

Methodology · Statistics 2015-09-28 Sayantan Banerjee , Rehan Akbani , Veerabhadran Baladandayuthapani

Modeling multiple sampling densities within a hierarchical framework enables borrowing of information across samples. These density random effects can act as kernels in latent variable models to represent exchangeable subgroups or clusters.…

Methodology · Statistics 2026-05-19 Yuliang Xu , Kaixuan Luo , Li Ma

We propose novel Bayesian Dynamic Clustering Factor Models (BDCFM) for the analysis of multivariate longitudinal data. BDCFM combines factor models with hidden Markov models to concomitantly perform dimension reduction, clustering, and…

Methodology · Statistics 2025-05-28 Tsering Dolkar , Marco A. R. Ferreira , Hwasoo Shin , Allison N. Tegge