English
Related papers

Related papers: Non-parametric Bayesian modelling of digital gene …

200 papers

In this dissertation, we develop nonparametric Bayesian models for biomedical data analysis. In particular, we focus on inference for tumor heterogeneity and inference for missing data. First, we present a Bayesian feature allocation model…

Applications · Statistics 2019-09-23 Tianjian Zhou

This report explores the application of machine learning techniques on short timeseries gene expression data. Although standard machine learning algorithms work well on longer time-series', they often fail to find meaningful insights from…

Genomics · Quantitative Biology 2021-11-17 Akankshita Dash

Simultaneous analysis of gene expression data and genetic variants is highly of interest, especially when the number of gene expressions and genetic variants are both greater than the sample size. Association of both causal genes and…

Methodology · Statistics 2021-10-07 Morteza Amini

Parametric Bayesian modeling offers a powerful and flexible toolbox for machine learning. Yet the model, however detailed, may still be wrong, and this can make inferences untrustworthy. In this paper we introduce a new class of…

Methodology · Statistics 2026-04-03 Bohan Wu , Eli N. Weinstein , Sohrab Salehi , Yixin Wang , David M. Blei

Discovering causal genetic variants from large genetic association studies poses many difficult challenges. Assessing which genetic markers are involved in determining trait status is a computationally demanding task, especially in the…

Genomics · Quantitative Biology 2015-04-09 Andrew L. Beam , Alison Motsinger-Reif , Jon Doyle

We develop a Bayesian nonparametric extension of the popular Plackett-Luce choice model that can handle an infinite number of choice items. Our framework is based on the theory of random atomic measures, with the prior specified by a gamma…

Machine Learning · Statistics 2012-11-20 Francois Caron , Yee Whye Teh

In computational biology, gene expression datasets are characterized by very few individual samples compared to a large number of measurements per sample. Thus, it is appealing to merge these datasets in order to increase the number of…

Methodology · Statistics 2011-08-18 Meili Baragatti

Because of the decreasing cost and high digital resolution, next-generation sequencing (NGS) is expected to replace the traditional hybridization-based microarray technology. For genetics study, the first-step analysis of NGS data is often…

Applications · Statistics 2014-01-13 Zhigen Zhao , Wei Wang , Zhi Wei

The advances of next-generation sequencing technology have accelerated study of the microbiome and stimulated the high throughput profiling of metagenomes. The large volume of sequenced data has encouraged the rise of various studies for…

Methodology · Statistics 2019-04-30 Qiwei Li , Shuang Jiang , Andrew Y. Koh , Guanghua Xiao , Xiaowei Zhan

Network models provide a powerful framework for analysing single-cell count data, facilitating the characterisation of cellular identities, disease mechanisms, and developmental trajectories. However, uncertainty modeling in unsupervised…

Genomics · Quantitative Biology 2026-04-27 Shanshan Ren , Thomas E. Bartlett , Lina Gerontogianni , Swati Chandna

Despite remarkable performance in producing realistic samples, Generative Adversarial Networks (GANs) often produce low-quality samples near low-density regions of the data manifold, e.g., samples of minor groups. Many techniques have been…

Machine Learning · Computer Science 2021-10-28 Jinhee Lee , Haeri Kim , Youngkyu Hong , Hye Won Chung

Microarrays are made it possible to simultaneously monitor the expression profiles of thousands of genes under various experimental conditions. Identification of co-expressed genes and coherent patterns is the central goal in microarray or…

Computational Engineering, Finance, and Science · Computer Science 2013-07-15 T. Chandrasekhar , K. Thangavel , E. Elayaraja , E. N. Sathishkumar

High-throughput scientific studies involving no clear a'priori hypothesis are common. For example, a large-scale genomic study of a disease may examine thousands of genes without hypothesizing that any specific gene is responsible for the…

Methodology · Statistics 2012-03-02 Babak Shahbaba

Traditional clustering methods are limited when dealing with huge and heterogeneous groups of gene expression data, which motivates the development of bi-clustering methods. Bi-clustering methods are used to mine bi-clusters whose subsets…

Computer Vision and Pattern Recognition · Computer Science 2020-05-13 Kaijie Xu , Witold Pedrycz , Zhiwu Li , Yinghui Quan , Weike Nie

The Generative Adversarial Network (GAN) was recently introduced in the literature as a novel machine learning method for training generative models. It has many applications in statistics such as nonparametric clustering and nonparametric…

Machine Learning · Statistics 2023-06-26 Sehwan Kim , Qifan Song , Faming Liang

Gene regulatory networks are collections of genes that interact with one other and with other substances in the cell. By measuring gene expression over time using high-throughput technologies, it may be possible to reverse engineer, or…

Applications · Statistics 2011-09-08 Andrea Rau , Florence Jaffrézic , Jean-Louis Foulley , R. W. Doerge

The problem of joint estimation of multiple graphical models from high dimensional data has been studied in the statistics and machine learning literature, due to its importance in diverse fields including molecular biology, neuroscience…

Methodology · Statistics 2019-07-04 Peyman Jalali , Kshitij Khare , George Michailidis

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

Applications · Statistics 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

The two-phase sampling design is a cost-efficient way of collecting expensive covariate information on a judiciously selected subsample. It is natural to apply such a strategy for collecting genetic data in a subsample enriched for exposure…

Applications · Statistics 2013-05-27 Jaeil Ahn , Bhramar Mukherjee , Stephen B. Gruber , Malay Ghosh

Statistical analysis of DNA mixtures is known to pose computational challenges due to the enormous state space of possible DNA profiles. We propose a Bayesian network representation for genotypes, allowing computations to be performed…

Methodology · Statistics 2014-02-21 Therese Graversen , Steffen Lauritzen