English
Related papers

Related papers: Rate optimal multiple testing procedure in high-di…

200 papers

Many statistical methodologies for high-dimensional data assume the population is normal. Although a few multivariate normality tests have been proposed, to the best of our knowledge, none of them can properly control the type I error when…

Methodology · Statistics 2021-05-04 Hao Chen , Yin Xia

In variable selection, most existing screening methods focus on marginal effects and ignore dependence between covariates. To improve the performance of selection, we incorporate pairwise effects in covariates for screening and…

Methodology · Statistics 2019-02-12 Siliang Gong , Kai Zhang , Yufeng Liu

Bayesian nonparametric regression under a rescaled Gaussian process prior offers smoothness-adaptive function estimation with near minimax-optimal error rates. Hierarchical extensions of this approach, equipped with stochastic variable…

Statistics Theory · Mathematics 2020-12-15 Sheng Jiang , Surya T. Tokdar

In this article, we propose a new algorithm for supervised learning methods, by which one can both capture the non-linearity in data and also find the best subset model. To produce an enhanced subset of the original variables, an ideal…

Applications · Statistics 2017-01-23 Peyman Tavallali , Marianne Razavi , Sean Brady

We propose sequential multiple testing procedures which control the false discover rate (FDR) or the positive false discovery rate (pFDR) under arbitrary dependence between the data streams. This is accomplished by "optimizing" an upper…

Methodology · Statistics 2024-11-27 Michael Hankin , Jay Bartroff

Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to managing false positives across many simultaneous tests. Hypotheses rarely exist in isolation;…

Methodology · Statistics 2026-05-19 Binyamin Perets , Shie Mannor

We investigate asymptotically optimal multiple testing procedures for streams of sequential data in the context of prior information on the number of false null hypotheses ("signals"). We show that the "gap" and "gap-intersection"…

Methodology · Statistics 2020-05-04 Xinrui He , Jay Bartroff

High-dimensional linear regression has been thoroughly studied in the context of independent and identically distributed data. We propose to investigate high-dimensional regression models for independent but non-identically distributed…

Statistics Theory · Mathematics 2026-05-20 Jérémie Bigot , Issa-Mbenard Dabo , Camille Male

Differential privacy provides a rigorous framework for privacy-preserving data analysis. This paper proposes the first differentially private procedure for controlling the false discovery rate (FDR) in multiple hypothesis testing. Inspired…

Statistics Theory · Mathematics 2021-07-06 Cynthia Dwork , Weijie J. Su , Li Zhang

A new online multiple testing procedure is described in the context of anomaly detection, which controls the False Discovery Rate (FDR). An accurate anomaly detector must control the false positive rate at a prescribed level while keeping…

Methodology · Statistics 2024-12-17 Etienne Krönert , Alain Célisse , Dalila Hattab

This paper explores the multiple testing problem for sparse high-dimensional data with binary outcomes. We propose novel empirical Bayes multiple testing procedures based on a spike-and-slab posterior and then evaluate their performance in…

Statistics Theory · Mathematics 2025-06-16 Yu-Chien Bo Ning

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split…

Methodology · Statistics 2018-05-04 Rina Foygel Barber , Emmanuel J. Candes

Businesses frequently run online controlled experiments (i.e., A/B tests) to learn about the effect of an intervention on multiple business metrics. To account for multiple hypothesis testing, multiple metrics are commonly aggregated into a…

Methodology · Statistics 2026-01-22 Luke Hagar , Nathaniel T. Stevens

Testing composite null hypotheses arises in various applications, such as mediation and replicability analyses. The problem becomes more challenging in high-throughput experiments where tens of thousands of features are examined…

Methodology · Statistics 2025-04-29 Pengfei Lyu , Xianyang Zhang , Hongyuan Cao

We introduce a multiple testing procedure that controls the median of the proportion of false discoveries (FDP) in a flexible way. The procedure only requires a vector of p-values as input and is comparable to the Benjamini-Hochberg method,…

Methodology · Statistics 2024-03-14 Jesse Hemerik , Aldo Solari , Jelle J Goeman

This paper studies how insurers can chose which claims to investigate for fraud. Given a prediction model, typically only claims with the highest predicted propability of being fraudulent are investigated. We argue that this can lead to…

Machine Learning · Statistics 2025-09-24 Christos Revelas , Otilia Boldea , Bas J. M. Werker

Subset selection in multiple linear regression aims to choose a subset of candidate explanatory variables that tradeoff fitting error (explanatory power) and model complexity (number of variables selected). We build mathematical programming…

Machine Learning · Statistics 2020-09-04 Young Woong Park , Diego Klabjan

A resurgence of interest in multiple hypothesis testing has occurred in the last decade. Motivated by studies in genomics, microarrays, DNA sequencing, drug screening, clinical trials, bioassays, education and psychology, statisticians have…

Statistics Theory · Mathematics 2007-06-13 Arthur Cohen , Harold B. Sackrowitz

This paper considers the problem of variable selection allowing for parameter instability. It distinguishes between signal and pseudo-signal variables that are correlated with the target variable, and noise variables that are not, and…

Econometrics · Economics 2024-07-17 Alexander Chudik , M. Hashem Pesaran , Mahrad Sharifvaghefi

Penalized regression models are popularly used in high-dimensional data analysis to conduct variable selection and model fitting simultaneously. Whereas success has been widely reported in literature, their performances largely depend on…

Machine Learning · Statistics 2013-12-16 Wei Sun , Junhui Wang , Yixin Fang
‹ Prev 1 8 9 10 Next ›