中文
相关论文

相关论文: Knockoffs with Side Information

200 篇论文

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

统计方法学 · 统计学 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

We study the problem of compressing a source sequence in the presence of side-information that is related to the source via insertions, deletions and substitutions. We propose a simple algorithm to compress the source sequence when the…

信息论 · 计算机科学 2016-11-15 Nan Ma , Kannan Ramchandran , David Tse

Matched case-control studies are commonly employed in epidemiological research for their convenience and efficiency. Analysis of secondary outcomes can yield valuable insights into biological pathways and help identify genetic variants of…

统计方法学 · 统计学 2026-02-24 Shanshan Liu , Guoqing Diao

The increasing availability of passively observed data has yielded a growing methodological interest in "data fusion." These methods involve merging data from observational and experimental sources to draw causal conclusions -- and they…

统计方法学 · 统计学 2021-12-15 Evan Rosenman , Art B. Owen

Multiple hypothesis testing, a situation when we wish to consider many hypotheses, is a core problem in statistical inference that arises in almost every scientific field. In this setting, controlling the false discovery rate (FDR), which…

统计理论 · 数学 2019-03-19 Shiyun Chen , Shiva Kasiviswanathan

In the context of multiple hypotheses testing, the proportion $\pi_0$ of true null hypotheses in the pool of hypotheses to test often plays a crucial role, although it is generally unknown a priori. A testing procedure using an implicit or…

统计理论 · 数学 2009-02-17 Gilles Blanchard , Etienne Roquain

Adaptive multiple testing with covariates is an important research direction that has gained major attention in recent years. It has been widely recognized that leveraging side information provided by auxiliary covariates can improve the…

统计方法学 · 统计学 2022-10-04 Dennis Leung , Wenguang Sun

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

统计方法学 · 统计学 2010-12-21 Xu Han , Weijie Gu , Jianqing Fan

Out of the participants in a randomized experiment with anticipated heterogeneous treatment effects, is it possible to identify which subjects have a positive treatment effect? While subgroup analysis has received attention, claims about…

统计方法学 · 统计学 2024-05-14 Boyan Duan , Larry Wasserman , Aaditya Ramdas

This paper investigates the benefits of the side information on the universal compression of sequences from a mixture of $K$ parametric sources. The output sequence of the mixture source is chosen from the source $i \in \{1,\ldots ,K\}$…

信息论 · 计算机科学 2014-12-01 Ahmad Beirami , Liling Huang , Mohsen Sardari , Faramarz Fekri

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

统计方法学 · 统计学 2025-05-08 Marco Molinari , Magne Thoresen

In general, if there is one device A with the same performance as many devices B, it would be better to replace many devices with one device. In order to determine the number of devices that can be reduced, it is important to determine the…

信息论 · 计算机科学 2019-05-30 Kiminori Iriyama

Balancing influential covariates is crucial for valid treatment comparisons in clinical studies. While covariate-adaptive randomization is commonly used to achieve balance, its performance can be inadequate when the number of baseline…

统计方法学 · 统计学 2024-12-30 Ziqing Guo , Yang Liu , Lucy Xia

We provide an approach to exploratory data analysis in matched observational studies with a single intervention and multiple endpoints. In such settings, the researcher would like to explore evidence for actual treatment effects among these…

统计方法学 · 统计学 2025-12-10 Mengqi Lin , Colin Fogarty

This paper demonstrates the potential of statistical disclosure control for protecting the data used to train recommender systems. Specifically, we use a synthetic data generation approach to hide specific information in the user-item…

信息检索 · 计算机科学 2020-08-11 Manel Slokom , Martha Larson , Alan Hanjalic

There is an increasing concern that most current published research findings are false. The main cause seems to lie in the fundamental disconnection between theory and practice in data analysis. While the former typically relies on…

机器学习 · 统计学 2019-03-06 Amedeo Roberto Esposito , Michael Gastpar , Ibrahim Issa

In this paper, we propose a model-based clustering method (TVClust) that robustly incorporates noisy side information as soft-constraints and aims to seek a consensus between side information and the observed data. Our method is based on a…

机器学习 · 统计学 2015-11-03 Daniel Khashabi , John Wieting , Jeffrey Yufei Liu , Feng Liang

Vovk (2015) introduced cross-conformal prediction, a modification of split conformal designed to improve the width of prediction sets. The method, when trained with a miscoverage rate equal to $\alpha$ and $n \gg K$, ensures a marginal…

机器学习 · 统计学 2025-05-22 Matteo Gasparin , Aaditya Ramdas

Selecting relevant features associated with a given response variable is an important issue in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent…

统计方法学 · 统计学 2020-12-17 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

The False Discovery Rate (FDR) is a new statistical procedure to control the number of mistakes made when performing multiple hypothesis tests, i.e. when comparing many data against a given model hypothesis. The key advantage of FDR is that…