中文
相关论文

相关论文: Multilayer Knockoff Filter: Controlled variable se…

200 篇论文

The interactions of users and items in recommender system could be naturally modeled as a user-item bipartite graph. In recent years, we have witnessed an emerging research effort in exploring user-item graph for collaborative filtering…

机器学习 · 计算机科学 2019-11-26 Xiao Wang , Ruijia Wang , Chuan Shi , Guojie Song , Qingyong Li

Recent discoveries suggest that our gut microbiome plays an important role in our health and wellbeing. However, the gut microbiome data are intricate; for example, the microbial diversity in the gut makes the data high-dimensional. While…

统计方法学 · 统计学 2021-03-02 Fang Xie , Johannes Lederer

Natural disasters, such as hurricanes and typhoons, pose significant challenges to public safety and infrastructure. While government agencies rely on multi million dollar UAV systems for storm data collection and disaster response, smaller…

动力系统 · 数学 2025-09-17 Ahmed A. Elgohary , Benjamin Gwinnell , Josh Augustine

Deep neural networks (DNNs) are famous for their high prediction accuracy, but they are also known for their black-box nature and poor interpretability. We consider the problem of variable selection, that is, selecting the input variables…

机器学习 · 统计学 2019-09-18 Zixuan Song , Jun Li

Deep neural networks (DNN) have been used successfully in many scientific problems for their high prediction accuracy, but their application to genetic studies remains challenging due to their poor interpretability. In this paper, we…

机器学习 · 计算机科学 2021-10-01 Peyman H. Kassani , Fred Lu , Yann Le Guen , Zihuai He

In sponsored search advertising, keywords serve as an essential bridge linking advertisers, search users and search engines. Advertisers have to deal with a series of keyword decisions throughout the entire lifecycle of search advertising…

信息检索 · 计算机科学 2022-03-01 Yanwu Yang , Bernard J. Jansen , Yinghui Yang , Xunhua Guo , Daniel Zeng

Non-negative Matrix Factorization(NMF) algorithm can only be used to find low rank approximation of original non-negative data while Concept Factorization(CF) algorithm extends matrix factorization to single non-linear kernel space,…

机器学习 · 计算机科学 2024-10-29 Fei Li , Liang Du , Chaohong Ren

Spatio-temporal data sets are rapidly growing in size. For example, environmental variables are measured with ever-higher resolution by increasing numbers of automated sensors mounted on satellites and aircraft. Using such data, which are…

统计方法学 · 统计学 2019-11-14 Marcin Jurek , Matthias Katzfuss

Multi-label causal feature selection has attracted extensive attention in recent years. However, current methods primarily operate at the label level, treating each label variable as a monolithic entity and overlooking the fine-grained…

机器学习 · 计算机科学 2026-02-16 Wanfu Gao , Yanan Wang , Yonghao Li

Feature selection is a widely used dimension reduction technique to select feature subsets because of its interpretability. Many methods have been proposed and achieved good results, in which the relationships between adjacent data points…

机器学习 · 计算机科学 2020-06-01 Yan Min , Mao Ye , Liang Tian , Yulin Jian , Ce Zhu , Shangming Yang

Replicability is a lynchpin for credible discoveries. The partial conjunction (PC) p-value, which combines individual base p-values from multiple similar studies, can gauge whether a feature of interest exhibits replicated signals across…

统计方法学 · 统计学 2025-07-29 Ninh Tran , Dennis Leung

Human interpretability of deep neural networks' decisions is crucial, especially in domains where these directly affect human lives. Counterfactual explanations of already trained neural networks can be generated by perturbing input…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Oana-Iuliana Popescu , Maha Shadaydeh , Joachim Denzler

Developing feature selection algorithms that move beyond a pure correlational to a more causal analysis of observational data is an important problem in the sciences. Several algorithms attempt to do so by discovering the Markov blanket of…

机器学习 · 统计学 2014-05-06 Eric V. Strobl , Shyam Visweswaran

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

信号处理 · 电气工程与系统科学 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

Feature selection is vital for building effective predictive models, as it reduces dimensionality and emphasizes key features. However, current research often suffers from limited benchmarking and reliance on proprietary datasets. This…

机器学习 · 计算机科学 2025-07-16 Vanderson Rocha , Diego Kreutz , Gabriel Canto , Hendrio Bragança , Eduardo Feitosa

Genome-wide association studies (GWAS) have led to the discovery of numerous single nucleotide polymorphisms (SNPs) associated with various phenotypes and complex diseases. However, the identified genetic variants do not fully explain the…

统计方法学 · 统计学 2025-07-09 Dayeon Jung , Yewon Kim , Junyong Park

Feature selection involes identifying the most relevant subset of input features, with a view to improving generalization of predictive models by reducing overfitting. Directly searching for the most relevant combination of attributes is…

机器学习 · 计算机科学 2014-10-28 Jayadeva , Sanjit S. Batra , Siddharth Sabharwal

This paper provides two general classes of multiple decision functions where each member of the first class strongly controls the family-wise error rate (FWER), while each member of the second class strongly controls the false discovery…

统计理论 · 数学 2019-11-19 Edsel A. Pena , Joshua D. Habiger , Wensong Wu

This work develops a new multifidelity ensemble Kalman filter (MFEnKF) algorithm based on linear control variate framework. The approach allows for rigorous multifidelity extensions of the EnKF, where the uncertainty in coarser fidelities…

数值分析 · 数学 2020-07-03 Andrey A Popov , Changhong Mou , Traian Iliescu , Adrian Sandu

Barber and Cand\`es (2015) control of the FDR in feature selection relies on estimating the FDR by the number of knockoff wins +1 divided by the number of original wins. We study the necessity of the +1 in general settings.

统计方法学 · 统计学 2024-12-10 Andrew Rajchert , Uri Keich