English
Related papers

Related papers: Detecting mutations in mixed sample sequencing dat…

200 papers

This paper explores the multiple testing problem for sparse high-dimensional data with binary outcomes. We propose novel empirical Bayes multiple testing procedures based on a spike-and-slab posterior and then evaluate their performance in…

Statistics Theory · Mathematics 2025-06-16 Yu-Chien Bo Ning

Anomaly detection is the process of finding data points that deviate from a baseline. In a real-life setting, anomalies are usually unknown or extremely rare. Moreover, the detection must be accomplished in a timely manner or the risk of…

Machine Learning · Computer Science 2019-04-26 Mariem Ben Fadhel , Kofi Nyarko

Identifying differentially expressed genes from RNA sequencing data remains a challenging task because of the considerable uncertainties in parameter estimation and the small sample sizes in typical applications. Here we introduce Bayesian…

Applications · Statistics 2014-11-11 Matthias Katzfuss , Andreas Neudecker , Simon Anders , Julien Gagneur

Statistical dependence between hypotheses poses a significant challenge to the stability of large scale multiple hypotheses testing. Ignoring it often results in an unacceptably large spread in the false positive proportion even though the…

Methodology · Statistics 2018-10-15 Sairam Rayaprolu , Zhiyi Chi

Discrete Markov random fields are undirected graphical models that capture complex conditional dependencies between discrete variables. Conducting exact posterior inference in these models is often computationally challenging because…

Methodology · Statistics 2026-03-10 Giuseppe Arena , Maarten Marsman

DNA microarrays are a relatively new technology that can simultaneously measure the expression level of thousands of genes. They have become an important tool for a wide variety of biological experiments. One of the most common goals of DNA…

Methodology · Statistics 2013-07-02 Eric Bair

Bayesian entity resolution merges together multiple, noisy databases and returns the minimal collection of unique individuals represented, together with their true, latent record values. Bayesian methods allow flexible generative models…

Methodology · Statistics 2014-10-20 Tamara Broderick , Rebecca C. Steorts

Motivated by an increasing demand for models that can effectively describe features of complex multivariate time series, e.g. from sensor data in biomechanics, motion analysis, and sports science, we introduce a novel state-space modeling…

Methodology · Statistics 2025-06-23 Alice Giampino , Bernardo Nipoti , Marina Vannucci , Michele Guindani

The detection of similarities between long DNA and protein sequences is studied using concepts of statistical physics. It is shown that mutual similarities can be detected by sequence alignment methods only if their amount exceeds a…

Condensed Matter · Physics 2009-10-28 Terence Hwa , Michael Lassig

Modern cancer genomics datasets involve widely varying sizes and scales, measurement variables, and correlation structures. A fundamental analytical goal in these high-throughput studies is the development of general statistical techniques…

Methodology · Statistics 2022-04-12 Chiyu Gu , Veerabhadran Baladandayuthapani , Subharup Guha

Genome sequencing technology has improved significantly in few last years and resulted in abundance genetic data. Artificial intelligence has been employed to analyze genetic data in response to its sheer size and variability. Gene…

Genomics · Quantitative Biology 2023-03-17 Muhammad Anwari Leksono , Ayu Purwarianti

Computer Vision practitioners must thoroughly understand their model's performance, but conditional evaluation is complex and error-prone. In biometric verification, model performance over continuous covariates---real-number attributes of…

Machine Learning · Computer Science 2020-09-22 Mel McCurrie , Hamish Nicholson , Walter J. Scheirer , Samuel Anthony

Safe and reliable disclosure of information from confidential data is a challenging statistical problem. A common approach considers the generation of synthetic data, to be disclosed instead of the original data. Efficient approaches ought…

Methodology · Statistics 2024-03-04 Larissa N. A. Martins , Flávio B. Gonçalves , Thais P. Galletti

This paper describes a Bayesian statistical method for determining the genetic basis of a complex genetic trait. The method uses a sample of unrelated individuals classified into two groups, for example cases and controls. Each group is…

Genomics · Quantitative Biology 2008-02-21 Toby Johnson

We consider the problem of clustering nested or hierarchical data, where observations are grouped and there are both group-level and observation-level variables. In our motivating OneK1K dataset, observations consist of single-cell…

Methodology · Statistics 2026-04-14 Arhit Chakrabarti , Yang Ni , Yuchao Jiang , Bani K. Mallick

DNA methylation (DNAme) is a critical component of the epigenetic regulatory machinery and aberrations in DNAme patterns occur in many diseases, such as cancer. Mapping and understanding DNAme profiles offers considerable promise for…

We study the problem of detecting change points (CPs) that are characterized by a subset of dimensions in a multi-dimensional sequence. A method for detecting those CPs can be formulated as a two-stage method: one for selecting relevant…

Machine Learning · Statistics 2018-03-05 Yuta Umezu , Ichiro Takeuchi

Repetitive elements are important in genomic structures, functions and regulations, yet effective methods in precisely identifying repetitive elements in DNA sequences are not fully accessible, and the relationship between repetitive…

Genomics · Quantitative Biology 2016-08-03 Changchuan Yin

We provide an approach to exploratory data analysis in matched observational studies with a single intervention and multiple endpoints. In such settings, the researcher would like to explore evidence for actual treatment effects among these…

Methodology · Statistics 2025-12-10 Mengqi Lin , Colin Fogarty

Anomaly detection is the process of identifying cases, or groups of cases, that are in some way unusual and do not fit the general patterns present in the dataset. Numerous algorithms use discretization of numerical data in their detection…

Databases · Computer Science 2020-08-31 Ralph Foorthuis