English
Related papers

Related papers: Bayesian History Reconstruction of Complex Human G…

200 papers

High-throughput genetic and epigenetic data are often screened for associations with an observed phenotype. For example, one may wish to test hundreds of thousands of genetic variants, or DNA methylation sites, for an association with…

Methodology · Statistics 2017-10-20 Eric F. Lock , David B. Dunson

In recent years the importance of finding a meaningful pattern from huge datasets has become more challenging. Data miners try to adopt innovative methods to face this problem by applying feature selection methods. In this paper we propose…

Machine Learning · Computer Science 2014-03-11 Mehdi Naseriparsa , Amir-masoud Bidgoli , Touraj Varaee

We present a novel coupled two-way clustering approach to gene microarray data analysis. The main idea is to identify subsets of the genes and samples, such that when one of these is used to cluster the other, stable and significant…

Biological Physics · Physics 2009-11-06 G. Getz , E. Levine , E. Domany

Cancer progression and monotonic accumulation models were developed to discover dependencies in the irreversible acquisition of binary traits from cross-sectional data. They have been used in computational oncology and virology but also in…

Populations and Evolution · Quantitative Biology 2025-05-12 Ramon Diaz-Uriarte , Iain G. Johnston

Despite initial success, cancer therapies often fail due to the emergence of drug-resistant cells. In this study, we use a mathematical model to investigate how cancer evolves over time, specifically focusing on the state of the tumor when…

Probability · Mathematics 2023-08-01 Kevin Leder , Zicheng Wang

Multiple Sequences Alignment (MSA) of biological sequences is a fundamental problem in computational biology due to its critical significance in wide ranging applications including haplotype reconstruction, sequence homology, phylogenetic…

Distributed, Parallel, and Cluster Computing · Computer Science 2009-05-13 Fahad Saeed , Ashfaq Khokhar

In this paper, we introduce a novel and interpretable methodology to cluster subjects suffering from cancer, based on features extracted from their biopsies. Contrary to existing approaches, we propose here to capture complex patterns in…

Quantitative Methods · Quantitative Biology 2020-07-07 Yassine El Ouahidi , Matis Feller , Matthieu Talagas , Bastien Pasdeloup

A pedigree is a directed graph that describes how individuals are related through ancestry in a sexually-reproducing population. In this paper we explore the question of whether one can reconstruct a pedigree by just observing sequence data…

Populations and Evolution · Quantitative Biology 2007-06-19 Bhalchandra D. Thatte , Mike Steel

The multispecies coalescent process models the genealogical relationships of genes sampled from several species, enabling useful predictions about phenomena such as the discordance between the gene tree and the species phylogeny due to…

Populations and Evolution · Quantitative Biology 2020-12-11 Jakub Truszkowski , Celine Scornavacca , Fabio Pardi

We develop a Bayesian model for globular clusters composed of multiple stellar populations, extending earlier statistical models for open clusters composed of simple (single) stellar populations (vanDyk et al. 2009, Stein et al. 2013).…

Solar and Stellar Astrophysics · Physics 2016-07-27 D. C. Stenning , R. Wagner-Kaiser , E. Robinson , D. A. van Dyk , T. von Hippel , A. Sarajedini , N. Stein

Phylogenetic mixtures model the inhomogeneous molecular evolution commonly observed in data. The performance of phylogenetic reconstruction methods where the underlying data is generated by a mixture model has stimulated considerable recent…

Populations and Evolution · Quantitative Biology 2007-06-30 Frederick A. Matsen , Mike Steel

This paper presents a new modeling strategy for joint unsupervised analysis of multiple high-throughput biological studies. As in Multi-study Factor Analysis, our goals are to identify both common factors shared across studies and…

Applications · Statistics 2018-06-27 Roberta De Vito , Ruggero Bellio , Lorenzo Trippa , Giovanni Parmigiani

Gene and protein networks are very important to model complex large-scale systems in molecular biology. Inferring or reverseengineering such networks can be defined as the process of identifying gene/protein interactions from experimental…

Machine Learning · Computer Science 2017-03-10 Stefano Beretta , Mauro Castelli , Ivo Goncalves , Ivan Merelli , Daniele Ramazzotti

We introduce a new algorithm called {\sc Rec-Gen} for reconstructing the genealogy or \textit{pedigree} of an extant population purely from its genetic data. We justify our approach by giving a mathematical proof of the effectiveness of…

Data Structures and Algorithms · Computer Science 2020-05-11 Younhun Kim , Elchanan Mossel , Govind Ramnarayan , Paxton Turner

Collecting genomics data across multiple heterogeneous populations (e.g., across different cancer types) has the potential to improve our understanding of disease. Despite sequencing advances, though, resources often remain a constraint…

Methodology · Statistics 2024-03-05 Yunyi Shen , Lorenzo Masoero , Joshua G. Schraiber , Tamara Broderick

Rich data generating mechanisms are ubiquitous in this age of information and require complex statistical models to draw meaningful inference. While Bayesian analysis has seen enormous development in the last 30 years, benefitting from the…

Computation · Statistics 2022-10-20 Radu V. Craiu , Evgeny Levi

$n$-gram profiles have been successfully and widely used to analyse long sequences of potentially differing lengths for clustering or classification. Mainly, machine learning algorithms have been used for this purpose but, despite their…

Methodology · Statistics 2024-09-04 José A. Perusquía , Jim E. Griffin , Cristiano Villa

The variation in DNA copy number carries information on the modalities of genome evolution and misregulation of DNA replication in cancer cells; its study can be helpful to localize tumor suppressor genes, distinguish different populations…

Methodology · Statistics 2012-03-20 Zhongyang Zhang , Kenneth Lange , Chiara Sabatti

Data clustering, including problems such as finding network communities, can be put into a systematic framework by means of a Bayesian approach. The application of Bayesian approaches to real problems can be, however, quite challenging. In…

Data Analysis, Statistics and Probability · Physics 2008-09-28 Alexei Vazquez

Complex data features, such as unmodelled censored event times and variables with time-dependent effects, are common in cancer recurrence studies and pose challenges for Bayesian survival modelling. Current methodologies for predictive…

Methodology · Statistics 2026-01-12 Saku Suorsa , Aki Vehtari