English
Related papers

Related papers: The Terminating-Random Experiments Selector: Fast …

200 papers

Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting labels, their label quality is compromised by the unavoidable…

Machine Learning · Computer Science 2026-02-17 Huipeng Huang , Wenbo Liao , Huajun Xi , Hao Zeng , Mengchen Zhao , Hongxin Wei

In hypothesis testing, a false discovery occurs when a hypothesis is incorrectly rejected due to noise in the sample. When adaptively testing multiple hypotheses, the probability of a false discovery increases as more tests are performed.…

Machine Learning · Statistics 2020-10-22 Wanrong Zhang , Gautam Kamath , Rachel Cummings

Many model selection algorithms produce a path of fits specifying a sequence of increasingly complex models. Given such a sequence and the data used to produce them, we consider the problem of choosing the least complex model that is not…

Methodology · Statistics 2015-12-09 William Fithian , Jonathan Taylor , Robert Tibshirani , Ryan Tibshirani

Some effort has been undertaken over the last decade to provide conditions for the control of the false discovery rate by the linear step-up procedure (LSU) for testing $n$ hypotheses when test statistics are dependent. In this paper we…

Statistics Theory · Mathematics 2007-10-18 Helmut Finner , Thorsten Dickhaus , Markus Roters

The Model-X knockoffs is a practical methodology for variable selection, which stands out from other selection strategies since it allows for the control of the false discovery rate (FDR), relying on finite-sample guarantees. In this…

We propose a novel sparse tensor decomposition method, namely Tensor Truncated Power (TTP) method, that incorporates variable selection into the estimation of decomposition components. The sparsity is achieved via an efficient truncation…

Machine Learning · Statistics 2016-05-04 Will Wei Sun , Junwei Lu , Han Liu , Guang Cheng

We consider the problem of identifying whether findings replicate from one study of high dimension to another, when the primary study guides the selection of hypotheses to be examined in the follow-up study as well as when there is no…

Methodology · Statistics 2014-01-28 Marina Bogomolov , Ruth Heller

In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…

Methodology · Statistics 2023-10-10 Yushu Shi , Michael Martens

A scientist tests a continuous stream of hypotheses over time in the course of her investigation -- she does not test a predetermined, fixed number of hypotheses. The scientist wishes to make as many discoveries as possible while ensuring…

Methodology · Statistics 2023-11-14 Ziyu Xu , Aaditya Ramdas

In the context of multiple hypotheses testing, the proportion $\pi_0$ of true null hypotheses in the pool of hypotheses to test often plays a crucial role, although it is generally unknown a priori. A testing procedure using an implicit or…

Statistics Theory · Mathematics 2009-02-17 Gilles Blanchard , Etienne Roquain

Motivation: Target-decoy search (TDS) is currently the most popular strategy for estimating and controlling the false discovery rate (FDR) of peptide identifications in mass spectrometry-based shotgun proteomics. While this strategy is very…

Applications · Statistics 2015-01-06 Kun He , Yan Fu , Wen-Feng Zeng , Lan Luo , Hao Chi , Chao Liu , Lai-Yun Qing , Rui-Xiang Sun , Si-Min He

Identifying important features linked to a response variable is a fundamental task in various scientific domains. This article explores statistical inference for simulated Markov random fields in high-dimensional settings. We introduce a…

Machine Learning · Statistics 2024-01-23 Haoyu Wei , Xiaoyu Lei , Yixin Han , Huiming Zhang

Despite the popularity of the false discovery rate (FDR) as an error control metric for large-scale multiple testing, its close Bayesian counterpart the local false discovery rate (lfdr), defined as the posterior probability that a…

Methodology · Statistics 2023-09-22 Jake A. Soloff , Daniel Xiang , William Fithian

In many scenarios such as genome-wide association studies where dependences between variables commonly exist, it is often of interest to infer the interaction effects in the model. However, testing pairwise interactions among millions of…

Methodology · Statistics 2022-09-02 Jingyi Duan , Yang Ning , Xi Chen , Yong Chen

This paper presents a systematic framework for controlling false discovery rate in learning time-varying correlation networks from high-dimensional, non-linear, non-Gaussian and non-stationary time series with an increasing number of…

Methodology · Statistics 2025-12-16 Bufan Li , Lujia Bai , Weichi Wu

In this paper we develop a consistent variable selection procedure for GARCH-X models that identifies the truly relevant exogenous covariates influencing volatility dynamics. The proposed method is based on a multiple hypothesis testing…

Methodology · Statistics 2026-04-29 Adriano Zanin Zambom , Beck Saunders

Controlled feature selection aims to discover the features a response depends on while limiting the false discovery rate (FDR) to a predefined level. Recently, multiple deep-learning-based methods have been proposed to perform controlled…

Machine Learning · Statistics 2022-10-24 Derek Hansen , Brian Manzo , Jeffrey Regier

Recently, Barber and Cand\`es laid the theoretical foundation for a general framework for false discovery rate (FDR) control based on the notion of "knockoffs." A closely related FDR control methodology has long been employed in the…

Methodology · Statistics 2022-03-15 Dong Luo , Arya Ebadi , Yilun He , Kristen Emery , William Stafford Noble , Uri Keich

Many important tasks of large-scale recommender systems can be naturally cast as testing multiple linear forms for noisy matrix completion. These problems, however, present unique challenges because of the subtle bias-and-variance tradeoff…

Methodology · Statistics 2025-03-12 Wanteng Ma , Lilun Du , Dong Xia , Ming Yuan

The TREX is a recently introduced method for performing sparse high-dimensional regression. Despite its statistical promise as an alternative to the lasso, square-root lasso, and scaled lasso, the TREX is computationally challenging in that…

Machine Learning · Statistics 2021-04-01 Jacob Bien , Irina Gaynanova , Johannes Lederer , Christian Müller
‹ Prev 1 4 5 6 7 8 10 Next ›