中文
相关论文

相关论文: A Dirichlet Process Mixture Model for Clustering L…

200 篇论文

Clustering multivariate binary data is of interest in many scientific fields, including ecology, biomedicine, and social policy. Beyond heuristic clustering algorithms, such data can be modelled using multivariate Bernoulli mixture models.…

统计方法学 · 统计学 2026-04-24 Luisa Ferrari , Maria Franco Villoria , Garritt L. Page , Alex Laini

Cancer has become one of the most widespread diseases in the world. Specifically, breast cancer is diagnosed more often than any other type of cancer. However, breast cancer patients and their individual tumors are often unique. Identifying…

定量方法 · 定量生物学 2016-12-06 Chenzhe Qian

Model-based clustering methods for continuous data are well established and commonly used in a wide range of applications. However, model-based clustering methods for categorical data are less standard. Latent class analysis is a commonly…

统计方法学 · 统计学 2013-02-20 Isabella Gollini , Thomas Brendan Murphy

We propose a novel "tree-averaging" model that utilizes the ensemble of classification and regression trees (CART). Each constituent tree is estimated with a subset of similar data. We treat this grouping of subsets as Bayesian ensemble…

机器学习 · 统计学 2014-08-20 Leo L. Duan , John P. Clancy , Rhonda D. Szczesniak

Bi-clustering is a technique that allows for the simultaneous clustering of observations and features in a dataset. This technique is often used in bioinformatics, text mining, and time series analysis. An important advantage of…

统计计算 · 统计学 2023-02-09 Anastasiia Livochka , Ryan Browne , Sanjeena Subedi

Understanding treatment effect heterogeneity is vital for scientific and policy research. However, identifying and evaluating heterogeneous treatment effects pose significant challenges due to the typically unknown subgroup structure.…

统计方法学 · 统计学 2024-11-05 Kwangho Kim , Jisu Kim , Larry A. Wasserman , Edward H. Kennedy

The identification of disease-gene associations is instrumental in understanding the mechanisms of diseases and developing novel treatments. Besides identifying genes from RNA-Seq datasets, it is often necessary to identify gene clusters…

基因组学 · 定量生物学 2025-11-14 Jake R. Patock , Rinki Ratnapriya , Arko Barman

A model based clustering procedure for data of mixed type, clustMD, is developed using a latent variable model. It is proposed that a latent variable, following a mixture of Gaussian distributions, generates the observed data of mixed type.…

统计方法学 · 统计学 2015-11-06 Damien McParland , Isobel Claire Gormley

Quantitative analysis of large-scale data is often complicated by the presence of diverse subgroups, which reduce the accuracy of inferences they make on held-out data. To address the challenge of heterogeneous data analysis, we introduce…

机器学习 · 计算机科学 2021-09-01 Nazanin Alipourfard , Keith Burghardt , Kristina Lerman

Machine learning techniques have been widely used in natural language processing (NLP). However, as revealed by many recent studies, machine learning models often inherit and amplify the societal biases in data. Various metrics have been…

计算与语言 · 计算机科学 2020-10-07 Jieyu Zhao , Kai-Wei Chang

Due to the challenge posed by multi-source and heterogeneous data collected from diverse environments, causal relationships among features can exhibit variations influenced by different time spans, regions, or strategies. This diversity…

机器学习 · 计算机科学 2025-02-11 Lu Liu , Yang Tang , Kexuan Zhang , Qiyu Sun

Variation in the evolutionary process across the sites of nucleotide sequence alignments is well established, and is an increasingly pervasive feature of datasets composed of gene regions sampled from multiple loci and/or different genomes.…

种群与进化 · 定量生物学 2014-09-04 Brian R. Moore , Jim McGuire , Fredrik Ronquist , John P. Huelsenbeck

Clustering based on vibration responses, such as transmissibility functions (TFs), is promising in structural anomaly detection. However, most existing methods struggle to determine the optimal cluster number, handle high-dimensional…

机器学习 · 计算机科学 2025-10-21 Lin-Feng Mei , Wang-Ji Yan

In this work, we study the problem of partitioning a set of graphs into different groups such that the graphs in the same group are similar while the graphs in different groups are dissimilar. This problem was rarely studied previously,…

机器学习 · 计算机科学 2023-02-07 Jinyu Cai , Yi Han , Wenzhong Guo , Jicong Fan

Cancer is a number of related yet highly heterogeneous diseases. Correct identification of cancer subtypes is critical for clinical decisions. The advance in sequencing technologies has made it possible to study cancer based on abundant…

应用统计 · 统计学 2018-11-27 Xiaochun Chen , Honggang Wang , Donghui Yan

In binary-transaction data-mining, traditional frequent itemset mining often produces results which are not straightforward to interpret. To overcome this problem, probability models are often used to produce more compact and conclusive…

机器学习 · 计算机科学 2012-09-27 Ruefei He , Jonathan Shapiro

Biclustering is an unsupervised machine-learning approach aiming to cluster rows and columns simultaneously in a data matrix. Several biclustering algorithms have been proposed for handling numeric datasets. However, real-world data mining…

机器学习 · 计算机科学 2024-08-26 Adán José-García , Julie Jacques , Clément Chauvet , Vincent Sobanski , Clarisse Dhaenens

Motivation: Clustering techniques are routinely applied to identify patterns of co-expression in gene expression data. Co-regulation, and involvement of genes in similar cellular function, is subsequently inferred from the clusters which…

定量方法 · 定量生物学 2016-06-10 Patrick E. McSharry , Edmund J. Crampin

In this paper, we introduce a novel and interpretable methodology to cluster subjects suffering from cancer, based on features extracted from their biopsies. Contrary to existing approaches, we propose here to capture complex patterns in…

定量方法 · 定量生物学 2020-07-07 Yassine El Ouahidi , Matis Feller , Matthieu Talagas , Bastien Pasdeloup

It is often of interest to perform clustering on longitudinal data, yet it is difficult to formulate an intuitive model for which estimation is computationally feasible. We propose a model-based clustering method for clustering objects that…

统计方法学 · 统计学 2020-05-19 Daniel K. Sewell , Yuguo Chen , William Bernhard , Tracy Sulkin