English
Related papers

Related papers: Detecting differentially methylated regions in bis…

200 papers

An important phenomenon in high dimensional biological data is the presence of unobserved covariates that can have a significant impact on the measured response. When these factors are also correlated with the covariate(s) of interest (i.e.…

Methodology · Statistics 2018-02-02 Chris McKennan , Dan Nicolae

Accurate computational identification of DNA methylation is essential for understanding epigenetic regulation. Although deep learning excels in this binary classification task, its "black-box" nature impedes biological insight. We address…

Machine Learning · Computer Science 2026-02-27 Yi He , Yina Cao , Jixiu Zhai , Di Wang , Junxiao Kong , Tianchi Lu

Detecting and assessing statistical significance of differentially methylated regions (DMRs) is a fundamental task in methylome association studies. While the average differential methylation in different phenotype groups has been the…

Methodology · Statistics 2023-06-28 Xiaoyu Wang , Ming Yu , William Grady , Ziding Feng , Wei Sun , James Y Dai

Protein footprinting is a new methodology that is based on probing, typically with the use of mass spectrometry, of reactivity of different aminoacid residues to a modifying reagent. Data thus obtained allow one to make inferences about…

Molecular Networks · Quantitative Biology 2010-12-21 Pasquale Moio , Arman Kulyyassov , Damien Vertut , Luc Camoin , Erlan Ramankulov , Marc Lipinski , Vasily Ogryzko

Spatially-explicit estimates of population density, together with appropriate estimates of uncertainty, are required in many management contexts. Density Surface Models (DSMs) are a two-stage approach for estimating spatially-varying…

Methodology · Statistics 2021-02-25 Mark V Bravington , David L Miller , Sharon L Hedley

The Regression Discontinuity Design (RDD) is a quasi-experimental design that estimates the causal effect of a treatment when its assignment is defined by a threshold value for a continuous assignment variable. The RDD assumes that subjects…

Applications · Statistics 2020-03-27 Federico Ricciardi , Silvia Liverani , Gianluca Baio

Discrete biomarkers derived as cell densities or counts from tissue microarrays and immunostaining are widely used to study immune signatures in relation to survival outcomes in cancer. Although routinely collected, these signatures are not…

Next-generation sequencing technologies now constitute a method of choice to measure gene expression. Data to analyze are read counts, commonly modeled using Negative Binomial distributions. A relevant issue associated with this…

Methodology · Statistics 2014-11-10 Elisabetta Bonafede , Franck Picard , Stéphane Robin , Cinzia Viroli

We consider learning parameters of Binomial Hidden Markov Models, which may be used to model DNA methylation data. The standard algorithm for the problem is EM, which is computationally expensive for sequences of the scale of the mammalian…

Machine Learning · Computer Science 2018-02-08 Chicheng Zhang , Eran A. Mukamel , Kamalika Chaudhuri

There exist several endeavors proposing a new family of extended distributions using the beta-generating technique. This is a well-known mechanism in developing flexible distributions, by embedding the cumulative distribution function (cdf)…

Statistics Theory · Mathematics 2019-12-17 M. Arashi , A. Bekker , D. de Waal , S. Makgai

Cell-free DNA (cfDNA) analysis is a powerful, minimally invasive tool for monitoring disease progression, treatment response, and early detection. A major challenge, however, is accurately determining the tissue of origin, especially in…

Genomics · Quantitative Biology 2025-06-03 Keng-Jung Lee , Dharanya Sampath , Konstantinos Mavrommatis

Regional aggregates of health outcomes over delineated administrative units (e.g., states, counties, zip codes), or areal units, are widely used by epidemiologists to map mortality or incidence rates and capture geographic variation. To…

Methodology · Statistics 2022-05-03 Leiwen Gao , Sudipto Banerjee , Beate Ritz

Diffusion models excel at generative modeling (e.g., text-to-image) but sampling requires multiple denoising network passes, limiting practicality. Efforts such as progressive distillation or consistency distillation have shown promise by…

Machine Learning · Computer Science 2025-04-01 Risheek Garrepalli , Shweta Mahajan , Munawar Hayat , Fatih Porikli

Recent technological advances coupled with large sample sets have uncovered many factors underlying the genetic basis of traits and the predisposition to complex disease, but much is left to discover. A common thread to most genetic…

Applications · Statistics 2013-12-11 Andrew Crossett , Ann B. Lee , Lambertus Klei , Bernie Devlin , Kathryn Roeder

With the development of next generation sequencing technology, researchers have now been able to study the microbiome composition using direct sequencing, whose output are bacterial taxa counts for each microbiome sample. One goal of…

Applications · Statistics 2013-05-24 Jun Chen , Hongzhe Li

Anomalies are those deviating from the norm. Unsupervised anomaly detection often translates to identifying low density regions. Major problems arise when data is high-dimensional and mixed of discrete and continuous attributes. We propose…

Machine Learning · Computer Science 2016-10-21 Kien Do , Truyen Tran , Svetha Venkatesh

Several modern genomic technologies, such as DNA-Methylation arrays, measure spatially registered probes that number in the hundreds of thousands across multiplechromosomes. The measured probes are by themselves less interesting…

Applications · Statistics 2016-11-16 John Nagorski , Genevera I. Allen

Protein complexes involved in DNA mismatch repair appear to diffuse along dsDNA in order to locate a hemimethylated incision site via a dissociative mechanism. Here, we study the probability that these complexes locate a given target site…

Subcellular Processes · Quantitative Biology 2021-05-12 Kyle Crocker , James London , Andrés Medina , Richard Fishel , Ralf Bundschuh

We propose a resampling-based fast variable selection technique for detecting relevant single nucleotide polymorphisms (SNP) in a multi-marker mixed effect model. Due to computational complexity, current practice primarily involves testing…

Applications · Statistics 2025-04-30 Subhabrata Majumdar , Saonli Basu , Matt McGue , Snigdhansu Chatterjee

High-throughput genetic and epigenetic data are often screened for associations with an observed phenotype. For example, one may wish to test hundreds of thousands of genetic variants, or DNA methylation sites, for an association with…

Methodology · Statistics 2017-10-20 Eric F. Lock , David B. Dunson