中文
相关论文

相关论文: Knockoffs with Side Information

200 篇论文

Despite the popularity of feature importance (FI) measures in interpretable machine learning, the statistical adequacy of these methods is rarely discussed. From a statistical perspective, a major distinction is between analyzing a…

机器学习 · 统计学 2023-05-03 Kristin Blesch , David S. Watson , Marvin N. Wright

Computer vision tasks often have side information available that is helpful to solve the task. For example, for crowd counting, the camera perspective (e.g., camera angle and height) gives a clue about the appearance and scale of people in…

计算机视觉与模式识别 · 计算机科学 2016-11-22 Di Kang , Debarun Dhar , Antoni B. Chan

Recommender systems have become an essential tool to help resolve the information overload problem in recent decades. Traditional recommender systems, however, suffer from data sparsity and cold start problems. To address these issues, a…

信息检索 · 计算机科学 2020-07-21 Zhu Sun , Qing Guo , Jie Yang , Hui Fang , Guibing Guo , Jie Zhang , Robin Burke

Top-$N$ recommender systems typically utilize side information to address the problem of data sparsity. As nowadays side information is growing towards high dimensionality, the performances of existing methods deteriorate in terms of both…

信息检索 · 计算机科学 2017-05-17 Yifan Chen , Xiang Zhao

Recently, Barber and Cand\`es laid the theoretical foundation for a general framework for false discovery rate (FDR) control based on the notion of "knockoffs." A closely related FDR control methodology has long been employed in the…

统计方法学 · 统计学 2022-03-15 Dong Luo , Arya Ebadi , Yilun He , Kristen Emery , William Stafford Noble , Uri Keich

One challenge in exploratory association studies using observational data is that the associations between the predictors and the outcome are potentially weak and rare, and the candidate predictors have complex correlation structures. False…

统计方法学 · 统计学 2025-01-30 Runqiu Wang , Ran Dai , Hongying Dai , Evan French , Cheng Zheng

Competition-based approach to controlling the false discovery rate (FDR) recently rose to prominence when, generalizing it to sequential hypothesis testing, Barber and Cand\`es used it as part of their knockoff-filter. Control of the FDR…

统计方法学 · 统计学 2023-02-24 Arya Ebadi , Dong Luo , Jack Freestone , William Stafford Noble , Uri Keich

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

统计理论 · 数学 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

Predictive modeling often uses black box machine learning methods, such as deep neural networks, to achieve state-of-the-art performance. In scientific domains, the scientist often wishes to discover which features are actually important…

机器学习 · 统计学 2020-08-03 Mukund Sudarshan , Wesley Tansey , Rajesh Ranganath

Session-based recommendation is gaining increasing attention due to its practical value in predicting the intents of anonymous users based on limited behaviors. Emerging efforts incorporate various side information to alleviate inherent…

信息检索 · 计算机科学 2025-05-20 Xiaokun Zhang , Bo Xu , Chenliang Li , Bowei He , Hongfei Lin , Chen Ma , Fenglong Ma

This article presents a Conformalized Locally Adaptive Weighting (CLAW) approach to multiple testing with side information. The proposed method employs innovative data-driven strategies to construct pairwise exchangeable scores, which are…

统计方法学 · 统计学 2025-02-28 Zinan Zhao , Wenguang Sun

We propose new methods to obtain simultaneous false discovery proportion bounds for knockoff-based approaches. We first investigate an approach based on Janson and Su's $k$-familywise error rate control method and interpolation. We then…

统计方法学 · 统计学 2024-02-27 Jinzhou Li , Marloes H. Maathuis , Jelle J. Goeman

Feature selection prepares the AI-readiness of data by eliminating redundant features. Prior research falls into two primary categories: i) Supervised Feature Selection, which identifies the optimal feature subset based on their relevance…

机器学习 · 计算机科学 2024-03-08 Xinyuan Wang , Dongjie Wang , Wangyang Ying , Rui Xie , Haifeng Chen , Yanjie Fu

Controlling false discovery rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring possible…

机器学习 · 统计学 2021-01-26 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

We describe a series of algorithms that efficiently implement Gaussian model-X knockoffs to control the false discovery rate on large scale feature selection problems. Identifying the knockoff distribution requires solving a large scale…

机器学习 · 计算机科学 2020-06-17 Armin Askari , Quentin Rebjock , Alexandre d'Aspremont , Laurent El Ghaoui

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of the annotators, some…

机器学习 · 统计学 2016-06-17 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao

Background: Adaptive interventions provide a guide for using ongoing information about individuals to decide whether and how to modify the type, amount, delivery modality, or timing of treatment, to improve intervention effectiveness while…

统计方法学 · 统计学 2026-03-12 John J. Dziak , Inbal Nahum-Shani

The traditional framework for feature selection treats all features as costing the same amount. However, in reality, a scientist often has considerable discretion regarding which variables to measure, and the decision involves a tradeoff…

统计方法学 · 统计学 2023-02-14 Guo Yu , Daniela Witten , Jacob Bien

Large-scale assessment data typically include numerous categorical variables, often affected by missing values. Motivated by the challenges arising in this framework, we extend the knockoffs method for selecting predictors to settings with…

统计方法学 · 统计学 2026-05-13 Silvia Bacci , Emanuela Dreassi , Leonardo Grilli , Carla Rampichini

Statistical hypothesis testing serves as statistical evidence for scientific innovation. However, if the reported results are intentionally biased, hypothesis testing no longer controls the rate of false discovery. In particular, we study…

统计方法学 · 统计学 2018-10-12 Junpei Komiyama , Takanori Maehara