English
Related papers

Related papers: Mixed membership analysis of genome-wide expressio…

200 papers

Heterogeneity is a fundamental characteristic of cancer. To accommodate heterogeneity, subgroup identification has been extensively studied and broadly categorized into unsupervised and supervised analysis. Compared to unsupervised…

Methodology · Statistics 2026-02-25 Xing Qin , Xu Liu , Shuangge Ma , Mengyun Wu

A number of statistical models have been successfully developed for the analysis of high-throughput data from a single source, but few methods are available for integrating data from different sources. Here we focus on integrating gene…

In language processing, training data with extremely large variance may lead to difficulty in the language model's convergence. It is difficult for the network parameters to adapt sentences with largely varied semantics or grammatical…

Computation and Language · Computer Science 2022-05-26 Yunhao Yang , Zhaokun Xue

Contagion effect refers to the causal effect of peers' behavior on the outcome of an individual in social networks. Contagion can be confounded due to latent homophily which makes contagion effect estimation very hard: nodes in a homophilic…

Machine Learning · Computer Science 2023-10-19 Zahra Fatemi , Elena Zheleva

We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for…

Machine Learning · Computer Science 2017-10-18 Romain Lopez , Jeffrey Regier , Michael Cole , Michael Jordan , Nir Yosef

Integrative network modeling of data arising from multiple genomic platforms provides insight into the holistic picture of the interactive system, as well as the flow of information across many disease domains including cancer. The basic…

Methodology · Statistics 2020-02-18 Min Jin Ha , Francesco Stingo , Veerabhadran Baladandayuthapani

It is very challenging to select informative features from tens of thousands of measured features in high-throughput data analysis. Recently, several parametric/regression models have been developed utilizing the gene network information to…

Applications · Statistics 2014-08-01 Yize Zhao , Jian Kang , Tianwei Yu

Unbiased, label-free proteomics is becoming a powerful technique for measuring protein expression in almost any biological sample. The output of these measurements after preprocessing is a collection of features and their associated…

The advent of Scientific Machine Learning has heralded a transformative era in scientific discovery, driving progress across diverse domains. Central to this progress is uncovering scientific laws from experimental data through symbolic…

Methodology · Statistics 2025-09-25 Somjit Roy , Pritam Dey , Debdeep Pati , Bani K. Mallick

In many statistical problems, a more coarse-grained model may be suitable for population-level behaviour, whereas a more detailed model is appropriate for accurate modelling of individual behaviour. This raises the question of how to…

Machine Learning · Statistics 2015-11-02 Mingjun Zhong , Nigel Goddard , Charles Sutton

Real-life statistical samples are often plagued by selection bias, which complicates drawing conclusions about the general population. When learning causal relationships between the variables is of interest, the sample may be assumed to be…

Statistics Theory · Mathematics 2018-11-15 Angelos P. Armen , Robin J. Evans

Exposure to diverse non-genetic factors, known as the exposome, is a critical determinant of health outcomes. However, analyzing the exposome presents significant methodological challenges, including: high collinearity among exposures, the…

Methodology · Statistics 2025-10-10 Matteo Amestoy , Mark van de Wiel , Jeroen Lakerveld , Wessel van Wieringen

Probabilistic graphical models (PGMs) are powerful tools for representing statistical dependencies through graphs in high-dimensional systems. However, they are limited to pairwise interactions. In this work, we propose the simplicial…

Machine Learning · Statistics 2025-10-16 Lorenzo Marinucci , Gabriele D'Acunto , Paolo Di Lorenzo , Sergio Barbarossa

The advances of next-generation sequencing technology have accelerated study of the microbiome and stimulated the high throughput profiling of metagenomes. The large volume of sequenced data has encouraged the rise of various studies for…

Methodology · Statistics 2019-04-30 Qiwei Li , Shuang Jiang , Andrew Y. Koh , Guanghua Xiao , Xiaowei Zhan

In ecology, the description of species composition and biodiversity calls for statistical methods that involve estimating features of interest in unobserved samples based on an observed one. In the last decade, the Bayesian nonparametrics…

Methodology · Statistics 2026-04-28 Alessandro Colombi , Raffaele Argiento , Federico Camerlenghi , Lucia Paci

Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. In this work, we…

Machine Learning · Computer Science 2026-04-07 Andrew Nam , Declan Campbell , Thomas Griffiths , Jonathan Cohen , Sarah-Jane Leslie

One important problem in genome science is to determine sets of co-regulated genes based on measurements of gene expression levels across samples, where the quantification of expression levels includes substantial technical and biological…

Applications · Statistics 2013-10-18 Chuan Gao , Christopher D Brown , Barbara E Engelhardt

Bayesian networks are powerful statistical models to study the probabilistic relationships among set random variables with major applications in disease modeling and prediction. Here, we propose a continuous time Bayesian network with…

Machine Learning · Computer Science 2021-07-16 Syed Hasib Akhter Faruqui , Adel Alaeddini , Jing Wang , Carlos A. Jaramillo

Learning the structure of Bayesian networks from data provides insights into underlying processes and the causal relationships that generate the data, but its usefulness depends on the homogeneity of the data population, a condition often…

Estimating conditional independence graphs from high-dimensional Gaussian data is challenging because methods must detect relevant edges while rigorously controlling statistical errors. We propose a Bayesian framework based on a prior…

Methodology · Statistics 2026-04-21 Roland B. Sogan , Tabea Rebafka , Fanny Villers