English
Related papers

Related papers: Selecting ChIP-seq Normalization Methods from the …

200 papers

We investigate saddlepoint approximations applied to the score test statistic in genome-wide association studies with binary phenotypes. The inaccuracy in the normal approximation of the score test statistic increases with increasing sample…

In cell differentiation, a cell of a less specialized type becomes one of a more specialized type, even though all cells have the same genome. Transcription factors and epigenetic marks like histone modifications can play a significant role…

Quantitative Methods · Quantitative Biology 2013-07-31 Nishanth Ulhas Nair , Yu Lin , Philipp Bucher , Bernard M. E. Moret

Two-sample inference for the difference of population means typically relies upon a Central Limit Theorem approximation. When data are drawn from a Negative Binomial distribution, previous work of Shilane et al. (2010) showed that a Normal…

Methodology · Statistics 2012-03-06 David Shilane , Derek Bean

Disjoint sampling is critical for rigorous and unbiased evaluation of state-of-the-art (SOTA) models. When training, validation, and test sets overlap or share data, it introduces a bias that inflates performance metrics and prevents…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Muhammad Ahmad , Manuel Mazzara , Salvatore Distifano

Histopathology-based survival modelling has two major hurdles. Firstly, a well-performing survival model has minimal clinical application if it does not contribute to the stratification of a cancer patient cohort into different risk groups,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-12 Hassan Muhammad , Chensu Xie , Carlie S. Sigel , Michael Doukas , Lindsay Alpert , William R. Jarnagin , Amber Simpson , Thomas J. Fuchs

Motivation: Researchers need a rich trove of genomic datasets that they can leverage to gain a better understanding of the genetic basis of the human genome and identify associations between phenotypes and specific parts of DNA. However,…

Cryptography and Security · Computer Science 2021-06-10 Nour Almadhoun Alserr , Gulce Kale , Onur Mutlu , Oznur Tastan , Erman Ayday

Gene transcription mediated by RNA polymerase II (pol-II) is a key step in gene expression. The dynamics of pol-II moving along the transcribed region influence the rate and timing of gene expression. In this work we present a probabilistic…

Although a vast body of literature relates to image segmentation methods that use deep neural networks (DNNs), less attention has been paid to assessing the statistical reliability of segmentation results. In this study, we interpret the…

Machine Learning · Statistics 2022-12-15 Vo Nguyen Le Duy , Shogo Iwazaki , Ichiro Takeuchi

Pooling genome-wide association studies of multiple related traits can substantially increase power for detecting genetic variants with pleiotropic effects. ASSET, which exhaustively searches all subsets of studies for association signals,…

Methodology · Statistics 2026-04-28 Samuel Anyaso-Samuel , Thong Luong , Fei Qin , Jiyeon Choi , Kai Yu , Paul S. Albert , Jianxin Shi

Effective disinfection is essential for maintaining water quality standards in distribution networks. Chlorination, as the most used technique, ensures safe water by maintaining sufficient chlorine residuals but also leads to the formation…

Systems and Control · Electrical Eng. & Systems 2024-09-16 Salma M. Elsherif , Ahmad F. Taha , Ahmed A. Abokifa

Statistically resolving the underlying haplotype pair for a genotype measurement is an important intermediate step in gene mapping studies, and has received much attention recently. Consequently, a variety of methods for this problem have…

Machine Learning · Computer Science 2007-10-29 Matti Kääriäinen , Niels Landwehr , Sampsa Lappalainen , Taneli Mielikäinen

Current popular methods in literature of RNA sequencing normalisation do not account for gene length when compared across samples, whilst adjusting for count biases in the data. This creates a gap in the normalisation as bigger genes in RNA…

Other Quantitative Biology · Quantitative Biology 2022-09-02 Hilbert Lam Yuen In , Robbe Pincket

We develop differentially private hypothesis testing methods for the small sample regime. Given a sample $\cal D$ from a categorical distribution $p$ over some domain $\Sigma$, an explicitly described distribution $q$ over $\Sigma$, some…

Data Structures and Algorithms · Computer Science 2017-06-08 Bryan Cai , Constantinos Daskalakis , Gautam Kamath

This paper introduces SelfMatch, a semi-supervised learning method that combines the power of contrastive self-supervised learning and consistency regularization. SelfMatch consists of two stages: (1) self-supervised pre-training based on…

Machine Learning · Computer Science 2021-01-19 Byoungjip Kim , Jinho Choo , Yeong-Dae Kwon , Seongho Joe , Seungjai Min , Youngjune Gwon

While the evaluation of explanations is an important step towards trustworthy models, it needs to be done carefully, and the employed metrics need to be well-understood. Specifically model randomization testing is often overestimated and…

In this paper, we study a class of two sample test statistics based on inter-point distances in the high dimensional and low sample size setting. Our test statistics include the well-known energy distance and maximum mean discrepancy with…

Methodology · Statistics 2020-04-13 Changbo Zhu , Xiaofeng Shao

We study the fundamental problems of identity testing (goodness of fit), and closeness testing (two sample test) of distributions over $k$ elements, under differential privacy. While the problems have a long history in statistics, finite…

Machine Learning · Computer Science 2017-11-01 Jayadev Acharya , Ziteng Sun , Huanyu Zhang

We study properties of two resampling scenarios: Conditional Randomisation and Conditional Permutation schemes, which are relevant for testing conditional independence of discrete random variables $X$ and $Y$ given a random variable $Z$.…

Statistics Theory · Mathematics 2023-04-14 Małgorzata Łazęcka , Bartosz Kołodziejek , Jan Mielniczuk

Previous divide-and-conquer segmentation analyses of DNA sequences do not provide a satisfactory stopping criterion for the recursion. This paper proposes that segmentation be considered as a model selection process. Using the tools in…

Biological Physics · Physics 2007-05-23 Wentian Li

As high-throughput sequencing has become common practice, the cost of sequencing large amounts of genetic data has been drastically reduced, leading to much larger data sets for analysis. One important task is to identify biological…

Methodology · Statistics 2014-10-14 Ciaran Evans , Johanna Hardin , Mark Huber , Daniel Stoebel , Garrett Wong