English
Related papers

Related papers: A General Framework for Multiple Testing via E-val…

200 papers

Replicability issues -- referring to the difficulty or failure of independent researchers to corroborate the results of published studies -- have hindered the meaningful progression of science and eroded public trust in scientific findings.…

Efforts to develop more efficient multiple hypothesis testing procedures for false discovery rate (FDR) control have focused on incorporating an estimate of the proportion of true null hypotheses (such procedures are called adaptive) or…

Methodology · Statistics 2017-02-13 Joshua D. Habiger

Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests to check whether those differences will generalise to online…

Information Retrieval · Computer Science 2025-07-23 David Otero , Javier Parapar , Álvaro Barreiro

Consider testing multiple hypotheses using tests that can only be evaluated by simulation, such as permutation tests or bootstrap tests. This article introduces MMCTest, a sequential algorithm which gives, with arbitrarily high probability,…

Methodology · Statistics 2018-10-17 Axel Gandy , Georg Hahn

In large-scale hypothesis testing, computing exact $p$-values or $e$-values is often resource-intensive, creating a need for budget-aware inferential methods. We propose a general framework for active hypothesis testing that leverages…

Methodology · Statistics 2026-04-09 Qi Kuang , Bowen Gang , Yin Xia

For multiple testing based on p-values with c\`{a}dl\`{a}g distribution functions, we propose an FDR procedure "BH+" with proven conservativeness. BH+ is at least as powerful as the BH procedure when they are applied to super-uniform…

Methodology · Statistics 2020-03-09 Xiongzhi Chen

We propose a novel multiple testing methodology for controlling the false discovery rate (FDR) in high-dimensional linear models that integrates model-X knockoff techniques with debiased penalized regression estimators. At the foundation of…

Methodology · Statistics 2026-03-17 Jinyuan Chang , Chenlong Li , Cheng Yong Tang , Zhengtian Zhu

Entity Matching (EM) is a core data cleaning task, aiming to identify different mentions of the same real-world entity. Active learning is one way to address the challenge of scarce labeled data in practice, by dynamically collecting the…

Databases · Computer Science 2020-03-31 Venkata Vamsikrishna Meduri , Lucian Popa , Prithviraj Sen , Mohamed Sarwat

The random coefficients model is an extension of the linear regression model that allows for unobserved heterogeneity in the population by modeling the regression coefficients as random variables. Given data from this model, the statistical…

Methodology · Statistics 2018-03-15 Fabian Dunker , Konstantin Eckle , Katharina Proksch , Johannes Schmidt-Hieber

Genome-wide association analysis has generated much discussion about how to preserve power to detect signals despite the detrimental effect of multiple testing on power. We develop a weighted multiple testing procedure that facilitates the…

Statistics Theory · Mathematics 2007-06-13 Kathryn Roeder , Bernie Devlin , Larry Wasserman

The analysis of large-scale datasets, especially in biomedical contexts, frequently involves a principled screening of multiple hypotheses. The celebrated two-group model jointly models the distribution of the test statistics with mixtures…

Methodology · Statistics 2023-03-10 Francesco Denti , Stefano Peluso , Michele Guindani , Antonietta Mira

Gaussian Mixture models (GMMs) are a powerful tool for clustering, classification and density estimation when clustering structures are embedded in the data. The presence of missing values can largely impact the GMMs estimation process,…

Machine Learning · Statistics 2020-06-05 Alessio Serafini , Thomas Brendan Murphy , Luca Scrucca

Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels).…

Machine Learning · Computer Science 2021-10-29 Maria Ulan , Welf Löwe , Morgan Ericsson , Anna Wingkvist

In big data applications such as healthcare data mining, due to privacy concerns, it is necessary to collect predictions from multiple information sources for the same instance, with raw features being discarded or withheld when aggregating…

Databases · Computer Science 2016-08-12 Chenwei Zhang , Sihong Xie , Yaliang Li , Jing Gao , Wei Fan , Philip S. Yu

We propose a unified class of calibration weighting methods based on weighted generalized entropy to handle missing at random (MAR) data with improved stability and efficiency. The proposed generalized entropy calibration (GEC) formulates…

Methodology · Statistics 2025-11-07 Yonghyun Kwon , Jae Kwang Kim , Yumou Qiu

This work proposes a novel rank-based scale two-sample testing method for univariate, distinct data when a subset of the data may be missing. Our approach is based on mathematically tight bounds of the Ansari-Bradley test statistic in the…

Methodology · Statistics 2025-09-25 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

The process comparing the empirical cumulative distribution function of the sample with a parametric estimate of the cumulative distribution function is known as the empirical process with estimated parameters and has been extensively…

Methodology · Statistics 2012-10-08 Ivan Kojadinovic , Jun Yan

Envelope methodology is succinctly pitched as a class of procedures for increasing efficiency in multivariate analyses without altering traditional objectives \citep[first sentence of page 1]{cook2018introduction}. This description is true…

Methodology · Statistics 2020-02-05 Daniel J. Eck

In neuroimaging, a large number of correlated tests are routinely performed to detect active voxels in single-subject experiments or to detect regions that differ between individuals belonging to different groups. In order to bound the…

Labeled data can be expensive to acquire in several application domains, including medical imaging, robotics, and computer vision. To efficiently train machine learning models under such high labeling costs, active learning (AL) judiciously…

Machine Learning · Computer Science 2022-06-13 Konstantinos D. Polyzos , Qin Lu , Georgios B. Giannakis