中文
相关论文

相关论文: Knoop: Practical Enhancement of Knockoff with Over…

200 篇论文

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so…

计算与语言 · 计算机科学 2018-06-20 Henry B. Moss , David S. Leslie , Paul Rayson

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

统计方法学 · 统计学 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

Model-based component-wise gradient boosting is a popular tool for data-driven variable selection. In order to improve its prediction and selection qualities even further, several modifications of the original algorithm have been developed,…

统计方法学 · 统计学 2023-02-28 Sophie Potts , Elisabeth Bergherr , Constantin Reinke , Colin Griesbach

Benkeser et al. demonstrate how adjustment for baseline covariates in randomized trials can meaningfully improve precision for a variety of outcome types. Their findings build on a long history, starting in 1932 with R.A. Fisher and…

统计方法学 · 统计学 2026-03-03 Laura B. Balzer , Erica Cai , Lucas Godoy Garraza , Pracheta Amaranath

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

统计方法学 · 统计学 2024-02-21 Zhaomeng Chen , Zihuai He , Benjamin B. Chu , Jiaqi Gu , Tim Morrison , Chiara Sabatti , Emmanuel Candès

Variable importance plays a pivotal role in interpretable machine learning as it helps measure the impact of factors on the output of the prediction model. Model agnostic methods based on the generation of "null" features via permutation…

The k-nearest neighbors (k-NN) classification rule has proven extremely successful in countless many computer vision applications. For example, image categorization often relies on uniform voting among the nearest prototypes in the space of…

计算机视觉与模式识别 · 计算机科学 2010-01-11 Paolo Piro , Richard Nock , Frank Nielsen , Michel Barlaud

As the size, complexity, and availability of data continues to grow, scientists are increasingly relying upon black-box learning algorithms that can often provide accurate predictions with minimal a priori model specifications. Tools like…

机器学习 · 统计学 2020-11-10 Lucas Mentch , Siyu Zhou

Decision trees and their ensembles are endowed with a rich set of diagnostic tools for ranking and screening variables in a predictive model. Despite the widespread use of tree based variable importance measures, pinning down their…

机器学习 · 统计学 2020-12-14 Jason M. Klusowski , Peter M. Tian

We present a novel framework for variable selection in Fr\'echet regression with responses in general metric spaces, a setting increasingly relevant for analyzing non-Euclidean data such as probability distributions and covariance matrices.…

统计理论 · 数学 2025-09-18 Haoyi Yang , Satarupa Bhattacharjee , Lingzhou Xue , Bing Li

Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel…

机器学习 · 统计学 2009-06-25 Genevera I. Allen

Recent advances in variational inference enable the modelling of highly structured joint distributions, but are limited in their capacity to scale to the high-dimensional setting of stochastic neural networks. This limitation motivates a…

Human interpretability of deep neural networks' decisions is crucial, especially in domains where these directly affect human lives. Counterfactual explanations of already trained neural networks can be generated by perturbing input…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Oana-Iuliana Popescu , Maha Shadaydeh , Joachim Denzler

We here introduce a novel classification approach adopted from the nonlinear model identification framework, which jointly addresses the feature selection and classifier design tasks. The classifier is constructed as a polynomial expansion…

机器学习 · 计算机科学 2016-07-29 Aida Brankovic , Alessandro Falsone , Maria Prandini , Luigi Piroddi

We apply the knockoff procedure to factor selection in finance. By building fake but realistic factors, this procedure makes it possible to control the fraction of false discovery in a given set of factors. To show its versatility, we apply…

统计金融 · 定量金融 2021-07-07 Damien Challet , Christian Bongiorno , Guillaume Pelletier

Standard approaches to tackle high-dimensional supervised classification problem often include variable selection and dimension reduction procedures. The novel methodology proposed in this paper combines clustering of variables and feature…

统计理论 · 数学 2018-11-07 Marie Chavent , Robin Genuer , Jerome Saracco

In modern day simulations of many-body systems much of the computational complexity is shifted to the identification of slowly changing molecular order parameters called collective variables (CV) or reaction coordinates. A vast array of…

统计力学 · 物理学 2016-04-27 Pratyush Tiwary , B. J. Berne

False discovery rate (FDR) controlling procedures provide important statistical guarantees for the replicability in signal identification based on multiple hypotheses testing. In many fields of study, FDR controlling procedures are used in…

统计方法学 · 统计学 2022-10-04 Ran Dai , Cheng Zheng

In this paper, we develop a general approach for probabilistic estimation and optimization. An explicit formula and a computational approach are established for controlling the reliability of probabilistic estimation based on a mixed…

统计理论 · 数学 2012-12-06 Xinjia Chen

In this paper we apply the previously introduced approximation method based on the ANOVA (analysis of variance) decomposition and Grouped Transformations to synthetic and real data. The advantage of this method is the interpretability of…

机器学习 · 统计学 2022-01-31 Daniel Potts , Michael Schmischke