中文
相关论文

相关论文: Robust inference with knockoffs

200 篇论文

This paper studies distribution-free inference in settings where the data set has a hierarchical structure -- for example, groups of observations, or repeated measurements. In such settings, standard notions of exchangeability may not hold.…

统计理论 · 数学 2025-08-05 Yonghoon Lee , Rina Foygel Barber , Rebecca Willett

Model selection is a central task in statistics, but standard methods are not robust in misspecified settings where the true data-generating process (DGP) is not in the set of candidate models. The key limitation is that existing methods --…

统计方法学 · 统计学 2026-03-10 Jongwoo Choi , Neil A. Spencer , Jeffrey W. Miller

Feature selection can be a crucial factor in obtaining robust and accurate predictions. Online feature selection models, however, operate under considerable restrictions; they need to efficiently extract salient input features based on a…

机器学习 · 计算机科学 2020-09-14 Johannes Haug , Martin Pawelczyk , Klaus Broelemann , Gjergji Kasneci

The model-X conditional randomization test is a generic framework for conditional independence testing, unlocking new possibilities to discover features that are conditionally associated with a response of interest while controlling type-I…

机器学习 · 计算机科学 2023-02-21 Shalev Shaer , Yaniv Romano

Conditional independence testing is an important problem, yet provably hard without assumptions. One of the assumptions that has become popular of late is called "model-X", where we assume we know the joint distribution of the covariates,…

统计方法学 · 统计学 2020-07-14 Eugene Katsevich , Aaditya Ramdas

We consider problems where many, somewhat redundant, hypotheses are tested and we are interested in reporting the most precise rejections, with false discovery rate (FDR) control. This is the case, for example, when researchers are…

统计方法学 · 统计学 2024-04-23 Paula Gablenz , Chiara Sabatti

Quantifying variable importance is essential for answering high-stakes questions in fields like genetics, public policy, and medicine. Current methods generally calculate variable importance for a given model trained on a given dataset.…

机器学习 · 计算机科学 2024-04-03 Jon Donnelly , Srikar Katta , Cynthia Rudin , Edward P. Browne

As the use of machine learning in high impact domains becomes widespread, the importance of evaluating safety has increased. An important aspect of this is evaluating how robust a model is to changes in setting or population, which…

机器学习 · 计算机科学 2021-03-16 Adarsh Subbaswamy , Roy Adams , Suchi Saria

Feature selection has remained a daunting challenge in machine learning and artificial intelligence, where increasingly complex, high-dimensional datasets demand principled strategies for isolating the most informative predictors. Despite…

机器学习 · 统计学 2025-12-02 Mousam Sinha , Tirtha Sarathi Ghosh , Ridam Pal

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

统计方法学 · 统计学 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

We consider a Bayesian approach to variable selection in the presence of high dimensional covariates based on a hierarchical model that places prior distributions on the regression coefficients as well as on the model space. We adopt the…

统计理论 · 数学 2014-07-28 Naveen Naidu Narisetty , Xuming He

Hypothesis testing in the linear regression model is a fundamental statistical problem. We consider linear regression in the high-dimensional regime where the number of parameters exceeds the number of samples ($p> n$). In order to make…

统计理论 · 数学 2019-09-24 Adel Javanmard , Jason D. Lee

Probabilistic model checking can provide formal guarantees on the behavior of stochastic models relating to a wide range of quantitative properties, such as runtime, energy consumption or cost. But decision making is typically with respect…

计算机科学中的逻辑 · 计算机科学 2024-03-19 Ingy Elsayed-Aly , David Parker , Lu Feng

Independence testing is a fundamental problem in statistical inference: given samples from a joint distribution $p$ over multiple random variables, the goal is to determine whether $p$ is a product distribution or is $\epsilon$-far from all…

机器学习 · 统计学 2026-03-06 Maryam Aliakbarpour , Alireza Azizi , Ria Stevens

In the high dimensional regression analysis when the number of predictors is much larger than the sample size, an important question is to select the important variable which are relevant to the response variable of interest. Variable…

统计方法学 · 统计学 2023-01-09 Pengsheng Ji , Zhigen Zhao

We give a finite-sample analysis of predictive inference procedures after model selection in regression with random design. The analysis is focused on a statistically challenging scenario where the number of potentially important…

统计理论 · 数学 2009-08-26 Hannes Leeb

We use a decision-theoretic framework to study the problem of forecasting discrete outcomes when the forecaster is unable to discriminate among a set of plausible forecast distributions because of partial identification or concerns about…

计量经济学 · 经济学 2020-12-18 Timothy Christensen , Hyungsik Roger Moon , Frank Schorfheide

Penalized regression methods are an attractive tool for high-dimensional data analysis, but their widespread adoption has been hampered by the difficulty of applying inferential tools. In particular, the question "How reliable is the…

统计理论 · 数学 2026-05-13 Patrick Breheny

A concept-based classifier can explain the decision process of a deep learning model by human-understandable concepts in image classification problems. However, sometimes concept-based explanations may cause false positives, which…

机器学习 · 计算机科学 2024-01-23 Kaiwen Xu , Kazuto Fukuchi , Youhei Akimoto , Jun Sakuma

We introduce a new family of one factor distributions for high-dimensional binary data. The model provides an explicit probability for each event, thus avoiding the numeric approximations often made by existing methods. Model interpretation…

统计方法学 · 统计学 2015-11-05 Matthieu Marbac , Mohammed Sedki