English
Related papers

Related papers: Marginal false discovery rates for penalized regre…

200 papers

Controlling False Discovery Rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring metrics…

Machine Learning · Statistics 2022-10-10 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

Identifying signals that replicate across multiple studies is essential for establishing robust scientific evidence, yet existing methods for high-dimensional replicability analysis either rely on restrictive modeling assumptions, are…

Methodology · Statistics 2026-03-05 Haochen Lei , Yan Li , Hongyuan Cao

The simultaneous analysis of many statistical tests is ubiquitous in applications. Perhaps the most popular error rate used for avoiding type one error inflation is the false discovery rate (FDR). However, most theoretical and software…

Computation · Statistics 2019-04-04 Guillermo Durand , Florian Junge , Sebastian Döhler , Etienne Roquain

As datasets grow richer, an important challenge is to leverage the full features in the data to maximize the number of useful discoveries while controlling for false positives. We address this problem in the context of multiple hypotheses…

Methodology · Statistics 2017-11-21 Fei Xia , Martin J. Zhang , James Zou , David Tse

Deep neural network (DNN) regression models are widely used in applications requiring state-of-the-art predictive accuracy. However, until recently there has been little work on accurate uncertainty quantification for predictions from such…

Methodology · Statistics 2020-09-07 Nadja Klein , David J. Nott , Michael Stanley Smith

This paper proposes new methodologies for conducting practical differentially private (DP) estimation and inference in high-dimensional linear regression. We first introduce a DP Bayesian Information Criterion (DP-BIC) for selecting the…

Methodology · Statistics 2026-04-13 Zhanrui Cai , Sai Li , Xintao Xia , Linjun Zhang

When estimating a regression model, we might have data where some labels are missing, or our data might be biased by a selection mechanism. When the response or selection mechanism is ignorable (i.e., independent of the response variable…

Statistics Theory · Mathematics 2023-08-22 Philip Boeken , Noud de Kroon , Mathijs de Jong , Joris M. Mooij , Onno Zoeter

Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to managing false positives across many simultaneous tests. Hypotheses rarely exist in isolation;…

Methodology · Statistics 2026-05-19 Binyamin Perets , Shie Mannor

In hypothesis testing, a false discovery occurs when a hypothesis is incorrectly rejected due to noise in the sample. When adaptively testing multiple hypotheses, the probability of a false discovery increases as more tests are performed.…

Machine Learning · Statistics 2020-10-22 Wanrong Zhang , Gautam Kamath , Rachel Cummings

In multiple hypothesis testing, it is well known that adaptive procedures can enhance power via incorporating information about the number of true nulls present. Under independence, we establish that two adaptive false discovery rate (FDR)…

Methodology · Statistics 2024-01-25 Dennis Leung , Ninh Tran

Histogram-based empirical Bayes methods developed for analyzing data for large numbers of genes, SNPs, or other biological features tend to have large biases when applied to data with a smaller number of features such as genes with…

Methodology · Statistics 2013-10-10 Marta Padilla , David R. Bickel

Testing composite null hypotheses arises in various applications, such as mediation and replicability analyses. The problem becomes more challenging in high-throughput experiments where tens of thousands of features are examined…

Methodology · Statistics 2025-04-29 Pengfei Lyu , Xianyang Zhang , Hongyuan Cao

Informed selection of drug candidates for laboratory experimentation provides an efficient means of identifying suitable anti-cancer treatments. The advancement of artificial intelligence has led to the development of computational models…

In genome-wide epigenetic studies, exposures (e.g., Single Nucleotide Polymorphisms) affect outcomes (e.g., gene expression) through intermediate variables such as DNA methylation. Mediation analysis offers a way to study these intermediate…

Methodology · Statistics 2025-12-08 Asmita Roy , Xianyang Zhang

In statistical disclosure control, the goal of data analysis is twofold: The released information must provide accurate and useful statistics about the underlying population of interest, while minimizing the potential for an individual…

Methodology · Statistics 2016-07-15 Jing Lei , Anne-Sophie Charest , Aleksandra Slavkovic , Adam Smith , Stephen Fienberg

Marginal expected shortfall is unquestionably one of the most popular systemic risk measures. Studying its extreme behaviour is particularly relevant for risk protection against severe global financial market downturns. In this context,…

Statistics Theory · Mathematics 2023-04-18 Simone A. Padoan , Stefano Rizzelli , Matteo Schiavone

Weak signal identification and inference are very important in the area of penalized model selection, yet they are under-developed and not well-studied. Existing inference procedures for penalized estimators are mainly focused on strong…

Methodology · Statistics 2016-11-16 Peibei Shi , Annie Qu

We propose a likelihood ratio based inferential framework for high dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high dimensional data analysis, including incomplete…

Machine Learning · Statistics 2015-11-24 Yang Ning , Tianqi Zhao , Han Liu

Inference of the marginal probability distribution is defined as the calculation of the probability of a subset of the variables and is relevant for handling missing data and hidden variables. While inference of the marginal probability…

Machine Learning · Statistics 2022-07-22 Fritz M. Bayer , Giusi Moffa , Niko Beerenwinkel , Jack Kuipers

The presence of confounding by high-dimensional variables complicates estimation of the average effect of a point treatment. On the one hand, it necessitates the use of variable selection strategies or more general data-adaptive…

Methodology · Statistics 2017-08-15 Vahe Avagyan , Stijn Vansteelandt