English
Related papers

Related papers: False discovery rate control with unknown null dis…

200 papers

We study a new framework for property testing of probability distributions, by considering distribution testing algorithms that have access to a conditional sampling oracle.* This is an oracle that takes as input a subset $S \subseteq [N]$…

Data Structures and Algorithms · Computer Science 2015-01-19 Clement Canonne , Dana Ron , Rocco A. Servedio

The Benjamini-Hochberg (BH) procedure is a celebrated method for multiple testing with false discovery rate (FDR) control. In this paper, we consider large-scale distributed networks where each node possesses a large number of p-values and…

Methodology · Statistics 2022-12-20 Mehrdad Pournaderi , Yu Xiang

We study the problem of Out-of-Distribution (OOD) detection, that is, detecting whether a learning algorithm's output can be trusted at inference time. While a number of tests for OOD detection have been proposed in prior work, a formal…

Machine Learning · Statistics 2023-09-19 Akshayaa Magesh , Venugopal V. Veeravalli , Anirban Roy , Susmit Jha

In the high dimensional regression analysis when the number of predictors is much larger than the sample size, an important question is to select the important variable which are relevant to the response variable of interest. Variable…

Methodology · Statistics 2023-01-09 Pengsheng Ji , Zhigen Zhao

In clinical and epidemiological research doubly truncated data often appear. This is the case, for instance, when the data registry is formed by interval sampling. Double truncation generally induces a sampling bias on the target variable,…

Methodology · Statistics 2023-01-11 Jacobo de Uña-Álvarez

Modern applications of conformal inference to multiple testing problems, such as outlier detection and candidate selection, often involve selecting test samples whose conformal p-values fall below a threshold. The quality of such methods is…

Methodology · Statistics 2026-05-21 Ziang Song , Ying Jin , Emmanuel J. Candès

We study the equivalence testing problem where the goal is to determine if the given two unknown distributions on $[n]$ are equal or $\epsilon$-far in the total variation distance in the conditional sampling model (CFGM, SICOMP16; CRS,…

Data Structures and Algorithms · Computer Science 2023-08-23 Diptarka Chakraborty , Sourav Chakraborty , Gunjan Kumar

We give oracle inequalities on procedures which combines quantization and variable selection via a weighted Lasso $k$-means type algorithm. The results are derived for a general family of weights, which can be tuned to size the influence of…

Statistics Theory · Mathematics 2016-07-07 Clément Levrard

Several classical methods exist for controlling the false discovery exceedance (FDX) for large scale multiple testing problems, among them the Lehmann-Romano procedure ([LR] below) and the Guo-Romano procedure ([GR] below). While these two…

Methodology · Statistics 2019-12-11 Sebastian Döhler , Etienne Roquain

There are many high dimensional function classes that have fast agnostic learning algorithms when assumptions on the distribution of examples can be made, such as Gaussianity or uniformity over the domain. But how can one be confident that…

Machine Learning · Computer Science 2022-11-22 Ronitt Rubinfeld , Arsen Vasilyan

As increasingly complex hypothesis-testing scenarios are considered in many scientific fields, analytic derivation of null distributions is often out of reach. To the rescue comes Monte Carlo testing, which may appear deceptively simple: as…

Methodology · Statistics 2015-04-13 Egil Ferkingstad , Lars Holden , Geir Kjetil Sandve

This article addresses a fundamental concern, first raised by Efron (2004), regarding the selection of null distributions in large-scale multiple testing. In modern data-intensive applications involving thousands or even millions of…

Methodology · Statistics 2025-09-05 Yang Tian , Zinan Zhao , Wenguang Sun

A new online multiple testing procedure is described in the context of anomaly detection, which controls the False Discovery Rate (FDR). An accurate anomaly detector must control the false positive rate at a prescribed level while keeping…

Methodology · Statistics 2024-12-17 Etienne Krönert , Alain Célisse , Dalila Hattab

Ensuring reliability is paramount in deep learning, particularly within the domain of medical imaging, where diagnostic decisions often hinge on model outputs. The capacity to separate out-of-distribution (OOD) samples has proven to be a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Anju Chhetri , Jari Korhonen , Prashnna Gyawali , Binod Bhattarai

The generalized linear models (GLM) have been widely used in practice to model non-Gaussian response variables. When the number of explanatory features is relatively large, scientific researchers are of interest to perform controlled…

Methodology · Statistics 2020-07-03 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

This paper considers Bayesian multiple testing under sparsity for polynomial-tailed distributions satisfying a monotone likelihood ratio property. Included in this class of distributions are the Student's t, the Pareto, and many other…

Statistics Theory · Mathematics 2016-07-29 Xueying Tang , Ke Li , Malay Ghosh

This paper discusses some problems possibly arising when approximating via Monte-Carlo simulations the distributions of goodness-of-fit test statistics based on the empirical distribution function. We argue that failing to re-estimate…

Data Analysis, Statistics and Probability · Physics 2008-04-01 Marco Capasso , Lucia Alessi , Matteo Barigozzi , Giorgio Fagiolo

Rao's spacing test is a widely used nonparametric method for assessing uniformity on the circle. However, its broader applicability in practical settings has been limited because the null distribution is not easily calculated. As a result,…

Methodology · Statistics 2026-02-05 Yoshiki Kinoshita , Aya Shinozaki , Toshinari Kamakura

Consider the problem of testing multiple null hypotheses. A classical approach to dealing with the multiplicity problem is to restrict attention to procedures that control the familywise error rate ($FWER$), the probability of even one…

Statistics Theory · Mathematics 2007-06-13 Joseph P. Romano , Azeem M. Shaikh

There is a well-known problem in Null Hypothesis Significance Testing: many statistically significant results fail to replicate in subsequent experiments. We show that this problem arises because standard `point-form null' significance…

Methodology · Statistics 2025-02-06 Fintan Costello , Paul Watts