English
Related papers

Related papers: Controlled Variable Selection from Summary Statist…

200 papers

Government statistical agencies often apply statistical disclosure limitation techniques to survey microdata to protect the confidentiality of respondents. There is a need for valid and practical ways to assess the protection provided. This…

Applications · Statistics 2010-11-15 Natalie Shlomo , Chris Skinner

False discovery rate (FDR) is a common way to control the number of false discoveries in multiple testing. There are a number of approaches available for controlling FDR. However, for functional test statistics, which are discretized into…

Methodology · Statistics 2024-12-03 Tomáš Mrkvička , Mari Myllymäki

Out of the participants in a randomized experiment with anticipated heterogeneous treatment effects, is it possible to identify which subjects have a positive treatment effect? While subgroup analysis has received attention, claims about…

Methodology · Statistics 2024-05-14 Boyan Duan , Larry Wasserman , Aaditya Ramdas

Algorithms that ensure reproducible findings from large-scale, high-dimensional data are pivotal in numerous signal processing applications. In recent years, multivariate false discovery rate (FDR) controlling methods have emerged,…

Methodology · Statistics 2024-01-31 Jasin Machkour , Michael Muma , Daniel P. Palomar

Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the…

Methodology · Statistics 2022-03-10 Xuebin Zhao , Hong Chen , Yingjie Wang , Weifu Li , Tieliang Gong , Yulong Wang , Feng Zheng

Omitted variables are one of the most important threats to the identification of causal effects. Several widely used methods assess the impact of omitted variables on empirical conclusions by comparing measures of selection on observables…

Econometrics · Economics 2026-02-05 Paul Diegert , Matthew A. Masten , Alexandre Poirier

One limitation of the most statistical/machine learning-based variable selection approaches is their inability to control the false selections. A recently introduced framework, model-x knockoffs, provides that to a wide range of models but…

Machine Learning · Statistics 2025-09-03 Deniz Koyuncu , Alex Gittens , Bülent Yener

Controlling the false discovery rate (FDR) is a popular approach to multiple testing, variable selection, and related problems of simultaneous inference. In many contemporary applications, models are not specified by discrete variables,…

Statistics Theory · Mathematics 2024-04-16 Mateo Díaz , Venkat Chandrasekaran

Knockoffs are a popular statistical framework that addresses the challenging problem of conditional variable selection in high-dimensional settings with statistical control. Such statistical control is essential for the reliability of…

Methodology · Statistics 2025-04-30 Alexandre Blain , Angel Reyero Lobo , Julia Linhart , Bertrand Thirion , Pierre Neuvial

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

Statistics Theory · Mathematics 2018-09-03 Irene Dowding , Stefan Haufe

Multivariate statistics are often available as well as necessary in hypothesis tests. We study how to use such statistics to control not only false discovery rate (FDR) but also positive FDR (pFDR) with good power. We show that FDR can be…

Statistics Theory · Mathematics 2008-05-21 Zhiyi Chi

An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , Amirata Ghorbani , James Zou

Random Forest has become one of the most popular tools for feature selection. Its ability to deal with high-dimensional data makes this algorithm especially useful for studies in neuroimaging and bioinformatics. Despite its popularity and…

Machine Learning · Computer Science 2014-10-13 Ender Konukoglu , Melanie Ganz

The collection and sharing of genomic data are becoming increasingly commonplace in research, clinical, and direct-to-consumer settings. The computational protocols typically adopted to protect individual privacy include sharing summary…

Cryptography and Security · Computer Science 2023-02-06 Rajagopal Venkatesaramani , Zhiyu Wan , Bradley A. Malin , Yevgeniy Vorobeychik

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

Methodology · Statistics 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

The mitigation of false positives is an important issue when conducting multiple hypothesis testing. The most popular paradigm for false positives mitigation in high-dimensional applications is via the control of the false discovery rate…

Methodology · Statistics 2018-07-17 Hien D. Nguyen , Yohan Yee , Geoffrey J. McLachlan , Jason P. Lerch

In high-dimensional data analysis, such as financial index tracking or biomedical applications, it is crucial to select the few relevant variables while maintaining control over the false discovery rate (FDR). In these applications, strong…

Portfolio Management · Quantitative Finance 2024-01-31 Jasin Machkour , Daniel P. Palomar , Michael Muma

As the number of parameters in Deep Neural Networks (DNNs) scales, the thirst for training data also increases. To save costs, it has become common for users and enterprises to delegate time-consuming data collection to third parties.…

Cryptography and Security · Computer Science 2023-10-17 Ziqiang Li , Pengfei Xia , Hong Sun , Yueqi Zeng , Wei Zhang , Bin Li

Estimating a causal effect from observational data can be biased if we do not control for self-selection. This selection is based on confounding variables that affect the treatment assignment and the outcome. Propensity score methods aim to…

Econometrics · Economics 2021-09-10 Daniel Jacob

Consider a case-control study in which we have a random sample, constructed in such a way that the proportion of cases in our sample is different from that in the general population---for instance, the sample is constructed to achieve a…

Methodology · Statistics 2019-01-01 Rina Foygel Barber , Emmanuel Candes