English
Related papers

Related papers: Mapping beyond diseases: Controlled variable selec…

200 papers

Models based on assumptions of multivariate regular variation and hidden regular variation provide ways to describe a broad range of extremal dependence structures when marginal distributions are heavy tailed. Multivariate regular variation…

Probability · Mathematics 2007-05-23 Janet E. Heffernan , Sidney I. Resnick

Multiple testing problems arising in modern scientific applications can involve simultaneously testing thousands or even millions of hypotheses, with relatively few true signals. In this paper, we consider the multiple testing problem where…

Methodology · Statistics 2016-06-28 Ang Li , Rina Foygel Barber

For randomized clinical trials where a single, primary, binary endpoint would require unfeasibly large sample sizes, composite endpoints are widely chosen as the primary endpoint. Despite being commonly used, composite endpoints entail…

Methodology · Statistics 2022-09-27 Marta Bofill Roig , Guadalupe Gómez Melis , Martin Posch , Franz Koenig

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

Methodology · Statistics 2025-05-08 Marco Molinari , Magne Thoresen

The positive false discovery rate (pFDR) is a useful overall measure of errors for multiple hypothesis testing, especially when the underlying goal is to attain one or more discoveries. Control of pFDR critically depends on how much…

Statistics Theory · Mathematics 2011-11-09 Zhiyi Chi

We propose one-at-a-time knockoffs (OATK), a new methodology for detecting important explanatory variables in linear regression models while controlling the false discovery rate (FDR). For each explanatory variable, OATK generates a…

Methodology · Statistics 2025-02-27 Charlie K. Guan , Zhimei Ren , Daniel W. Apley

In this paper we motivate the causal mechanisms behind sample selection induced collider bias (selection collider bias) that can cause Large Language Models (LLMs) to learn unconditional dependence between entities that are unconditionally…

Computation and Language · Computer Science 2022-09-14 Emily McMilin

This paper develops a model of \textit{identification design} and applies it to robust causal inference in microeconometrics. The decision maker observes the population distribution of signals generated by an information structure and ranks…

Theoretical Economics · Economics 2026-04-20 Maxwell Rosenthal

This research addresses the challenge of conducting interpretable causal inference between a binary treatment and its resulting outcome when not all confounders are known. Confounders are factors that have an influence on both the treatment…

Machine Learning · Computer Science 2023-10-24 Sohaib Kiani , Jared Barton , Jon Sushinsky , Lynda Heimbach , Bo Luo

The two-sample problem, which consists in testing whether independent samples on $\mathbb{R}^d$ are drawn from the same (unknown) distribution, finds applications in many areas. Its study in high-dimension is the subject of much attention,…

Statistics Theory · Mathematics 2023-02-09 Stephan Clémençon , Myrto Limnios , Nicolas Vayatis

Aleatoric and Epistemic uncertainty have achieved recent attention in the literature as different sources from which uncertainty can emerge in stochastic modeling. Epistemic being intrinsic or model based notions of uncertainty, and…

Methodology · Statistics 2025-08-15 Ryan Warnick

In many applications, we need to study a linear regression model that consists of a response variable and a large number of potential explanatory variables and determine which variables are truly associated with the response. In 2015,…

Methodology · Statistics 2019-07-23 Jiajie Chen , Anthony Hou , Thomas Y. Hou

In high dimensional analysis, effects of explanatory variables on responses sometimes rely on certain exposure variables, such as time or environmental factors. In this paper, to characterize the importance of each predictor, we utilize its…

Methodology · Statistics 2018-04-11 Yeqing Zhou , Jingyuan Liu , Zhihui Hao , Liping Zhu

Two recently introduced model based bias corrected estimators for proportion of true null hypotheses ($\pi_0$) under multiple hypotheses testing scenario have been restructured for exponentially distributed random observations available for…

Statistics Theory · Mathematics 2020-07-28 Aniket Biswas , Gaurangadeb Chattopadhyay , Aditya Chatterjee

Let $X=(X_1,\ldots,X_p)$ be a $p$-variate random vector and $F$ a fixed finite set. In a number of applications, mainly in genetics, it turns out that $X_i\in F$ for each $i=1,\ldots,p$. Despite the latter fact, to obtain a knockoff…

Statistics Theory · Mathematics 2024-10-15 Emanuela Dreassi , Luca Pratelli , Pietro Rigo

This article considers the problem of multiple hypothesis testing using $t$-tests. The observed data are assumed to be independently generated conditional on an underlying and unknown two-state hidden model. We propose an asymptotically…

Statistics Theory · Mathematics 2011-02-22 Hongyuan Cao , Michael R. Kosorok

Large-scale multiple testing with correlated and heavy-tailed data arises in a wide range of research areas from genomics, medical imaging to finance. Conventional methods for estimating the false discovery proportion (FDP) often ignore the…

Methodology · Statistics 2018-09-19 Jianqing Fan , Yuan Ke , Qiang Sun , Wen-Xin Zhou

Controlling the false discovery rate (FDR) is a powerful approach to multiple testing. In many applications, the tested hypotheses have an inherent hierarchical structure. In this paper, we focus on the fixed sequence structure where the…

Methodology · Statistics 2016-11-11 Gavin Lynch , Wenge Guo , Sanat K. Sarkar , Helmut Finner

Combining patient-level data from clinical trials can connect rare phenomena with clinical endpoints, but statistical techniques applied to a single trial may become problematical when trials are pooled. Estimating the hazard of a binary…

Consider the task of generating samples from a tilted distribution of a random vector whose underlying distribution is unknown, but samples from it are available. This finds applications in fields such as finance and climate science, and in…