English
Related papers

Related papers: Identifying statistical dependence in genomic sequ…

200 papers

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

Applications · Statistics 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

We consider a Bayesian functional data analysis for observations measured as extremely long sequences. Splitting the sequence into a number of small windows with manageable length, the windows may not be independent especially when they are…

Methodology · Statistics 2021-12-10 Suvo Chatterjee , Shrabanti Chowdhury , Duchwan Ryu , Sanjib Basu

DNA copy number and mRNA expression are widely used data types in cancer studies, which combined provide more insight than separately. Whereas in existing literature the form of the relationship between these two types of markers is fixed a…

A fundamental aspect of biological information processing is the ubiquity of sequence-function relationships -- functions that map the sequence of DNA, RNA, or protein to a biochemically relevant activity. Most sequence-function…

Quantitative Methods · Quantitative Biology 2016-03-23 Gurinder S. Atwal , Justin B. Kinney

Detecting variation in the evolutionary process along chromosomes is increasingly important as whole-genome data becomes more widely available. For example, factors such as incomplete lineage sorting, horizontal gene transfer, and…

Populations and Evolution · Quantitative Biology 2017-01-03 Elizabeth S. Allman , Laura S. Kubatko , John A. Rhodes

This paper considers the problem of matching fragment to organism using its complete genome. Our method is based on the probability measure representation of a genome. We first demonstrate that these probability measures can be modelled as…

Biological Physics · Physics 2009-11-07 V. V. Anh , K. S. Lau , Z. G. Yu

Mutual information (MI) is a fundamental measure of statistical dependence, with a myriad of applications to information theory, statistics, and machine learning. While it possesses many desirable structural properties, the estimation of…

Information Theory · Computer Science 2021-10-19 Ziv Goldfeld , Kristjan Greenewald

Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge…

Modern cancer genomics datasets involve widely varying sizes and scales, measurement variables, and correlation structures. A fundamental analytical goal in these high-throughput studies is the development of general statistical techniques…

Methodology · Statistics 2022-04-12 Chiyu Gu , Veerabhadran Baladandayuthapani , Subharup Guha

We are interested in the comparison of transcript boundaries from cells which originated in different environments. The goal is to assess whether this phenomenon, called differential splicing, is used to modify the transcription of the…

Applications · Statistics 2013-07-12 Alice Cleynen , Stéphane Robin

Alternative splicing is crucial in gene regulation, with significant implications in clinical settings and biotechnology. This review article compiles bioinformatics RNA-seq tools for investigating differential splicing; offering a detailed…

Genomics · Quantitative Biology 2024-09-10 Ben J Draper , Mark J Dunning , David C James

Determining the strength of non-linear statistical dependencies between two variables is a crucial matter in many research fields. The established measure for quantifying such relations is the mutual information. However, estimating mutual…

Data Analysis, Statistics and Probability · Physics 2019-07-24 Damián G. Hernández , Inés Samengo

This paper proposes a geometric estimator of dependency between a pair of multivariate samples. The proposed estimator of dependency is based on a randomly permuted geometric graph (the minimal spanning tree) over the two multivariate…

Machine Learning · Computer Science 2019-10-02 Salimeh Yasaei Sekeh , Alfred O. Hero

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding…

Machine Learning · Statistics 2016-10-04 Xin Gao , Raymond J. Carroll

A key task of data science is to identify relevant features linked to certain output variables that are supposed to be modeled or predicted. To obtain a small but meaningful model, it is important to find stochastically independent…

Methodology · Statistics 2021-12-23 Tim Breitenbach , Lauritz Rasbach , Chunguang Liang , Patrick Jahnke

Large-scale multiple testing tasks often exhibit dependence, and leveraging the dependence between individual tests is still one challenging and important problem in statistics. With recent advances in graphical models, it is feasible to…

Methodology · Statistics 2012-10-19 Jie Liu , Chunming Zhang , Catherine McCarty , Peggy Peissig , Elizabeth Burnside , David Page

Background: Coevolution within a protein family is often predicted using statistics that measure the degree of covariation between positions in the protein sequence. Mutual Information is a measure of dependence between two random variables…

Populations and Evolution · Quantitative Biology 2013-04-17 Russell J. Dickson , Gregory B. Gloor

Identification of essential genes is one of the ultimate goals of drug designs. Here we introduce an {\it in silico} method to select essential genes through the microarray assay. We construct a graph of genes, called the gene transcription…

Statistical Mechanics · Physics 2007-05-23 K. Rho , H. Jeong , B. Kahng

Transcriptomic data is a treasure-trove in modern molecular biology, as it offers a comprehensive viewpoint into the intricate nuances of gene expression dynamics underlying biological systems. This genetic information must be utilised to…

Molecular Networks · Quantitative Biology 2023-12-13 Vikram Singh , Vikram Singh

Motivation: Most existing methods for DNA sequence analysis rely on accurate sequences or genotypes. However, in applications of the next-generation sequencing (NGS), accurate genotypes may not be easily obtained (e.g. multi-sample…

Genomics · Quantitative Biology 2013-03-19 Heng Li