中文
相关论文

相关论文: A Generalized Knockoff Procedure for FDR Control i…

200 篇论文

Model-X knockoffs allows analysts to perform feature selection using almost any machine learning algorithm while still provably controlling the expected proportion of false discoveries. To apply model-X knockoffs, one must construct…

统计方法学 · 统计学 2021-06-30 Asher Spector , Lucas Janson

The False Discovery Rate (FDR) method has recently been described by Miller et al (2001), along with several examples of astrophysical applications. FDR is a new statistical procedure due to Benjamini and Hochberg (1995) for controlling the…

天体物理学 · 物理学 2009-11-07 A. M. Hopkins , C. J. Miller , A. J. Connolly , C. Genovese , R. C. Nichol , L. Wasserman

This paper proposes novel inferential procedures for discovering the network Granger causality in high-dimensional vector autoregressive models. In particular, we mainly offer two multiple testing procedures designed to control the false…

统计方法学 · 统计学 2024-11-14 Yoshimasa Uematsu , Takashi Yamagata

Barber and Cand\`es (2015) control of the FDR in feature selection relies on estimating the FDR by the number of knockoff wins +1 divided by the number of original wins. We study the necessity of the +1 in general settings.

统计方法学 · 统计学 2024-12-10 Andrew Rajchert , Uri Keich

In the multiple testing problem with independent tests, the classical linear step-up procedure controls the false discovery rate (FDR) at level $\pi_0\alpha$, where $\pi_0$ is the proportion of true null hypotheses and $\alpha$ is the…

统计方法学 · 统计学 2019-08-29 Peter MacDonald , Kun Liang , Arnold Janssen

Model-X knockoffs is a wrapper that transforms essentially any feature importance measure into a variable selection algorithm, which discovers true effects while rigorously controlling the expected fraction of false positives. A frequently…

统计方法学 · 统计学 2024-03-12 Stephen Bates , Emmanuel Candès , Lucas Janson , Wenshuo Wang

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of the annotators, some…

机器学习 · 统计学 2016-06-17 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao

In the context of high-dimensional Gaussian linear regression for ordered variables, we study the variable selection procedure via the minimization of the penalized least-squares criterion. We focus on model selection where the penalty…

统计理论 · 数学 2024-07-01 Perrine Lacroix , Marie-Laure Martin

Controlling False Discovery Rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring metrics…

机器学习 · 统计学 2022-10-10 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

Voxel-based multiple testing is widely used in neuroimaging data analysis. Traditional false discovery rate (FDR) control methods often ignore the spatial dependence among the voxel-based tests and thus suffer from substantial loss of…

机器学习 · 统计学 2024-05-06 Taehyo Kim , Hai Shu , Qiran Jia , Mony J. de Leon

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

统计方法学 · 统计学 2024-02-21 Zhaomeng Chen , Zihuai He , Benjamin B. Chu , Jiaqi Gu , Tim Morrison , Chiara Sabatti , Emmanuel Candès

Many approaches for multiple testing begin with the assumption that all tests in a given study should be combined into a global false-discovery-rate analysis. But this may be inappropriate for many of today's large-scale screening problems,…

统计方法学 · 统计学 2014-06-10 James G. Scott , Ryan C. Kelly , Matthew A. Smith , Pengcheng Zhou , Robert E. Kass

In this paper we introduce a novel approach for an important problem of break detection. Specifically, we are interested in detection of an abrupt change in the covariance structure of a high-dimensional random process -- a problem, which…

统计理论 · 数学 2020-07-30 Valeriy Avanesov , Nazar Buzun

High-dimensional longitudinal time series data is prevalent across various real-world applications. Many such applications can be modeled as regression problems with high-dimensional time series covariates. Deep learning has been a popular…

机器学习 · 统计学 2024-04-09 Wenxuan Zuo , Zifan Zhu , Yuxuan Du , Yi-Chun Yeh , Jed A. Fuhrman , Jinchi Lv , Yingying Fan , Fengzhu Sun

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

统计方法学 · 统计学 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

In this work we study an adaptive step-down procedure for testing $m$ hypotheses. It stems from the repeated use of the false discovery rate controlling the linear step-up procedure (sometimes called BH), and makes use of the critical…

统计理论 · 数学 2009-04-01 Yulia Gavrilov , Yoav Benjamini , Sanat K. Sarkar

Multiple testing with false discovery rate (FDR) control has been widely conducted in the ``discrete paradigm" where p-values have discrete and heterogeneous null distributions. However, in this scenario existing FDR procedures often lose…

统计方法学 · 统计学 2019-07-23 Xiongzhi Chen , R. W. Doerge , Sanat K. Sarkar

Competition-based FDR control has been commonly used for over a decade in the computational mass spectrometry community (Elias and Gygi, 2007). Recently, the approach has gained significant popularity in other fields after Barber and Candes…

统计方法学 · 统计学 2019-11-14 Kristen Emery , Syamand Hasam , William Stafford Noble , Uri Keich

We consider the variable selection problem, which seeks to identify important variables influencing a response $Y$ out of many candidate features $X_1, \ldots, X_p$. We wish to do so while offering finite-sample guarantees about the…

统计方法学 · 统计学 2019-02-12 Rina Foygel Barber , Emmanuel J. Candès , Richard J. Samworth

Consider the problem of simultaneously testing null hypotheses H_1,...,H_s. The usual approach to dealing with the multiplicity problem is to restrict attention to procedures that control the familywise error rate (FWER), the probability of…

统计理论 · 数学 2007-06-13 E. L. Lehmann , Joseph P. Romano