中文
相关论文

相关论文: Asymptotic Unbiasedness of the Permutation Importa…

200 篇论文

Throughout the last decade, random forests have established themselves as among the most accurate and popular supervised learning methods. While their black-box nature has made their mathematical analysis difficult, recent work has…

统计方法学 · 统计学 2019-12-10 Tim Coleman , Wei Peng , Lucas Mentch

This paper is about variable selection with the random forests algorithm in presence of correlated predictors. In high-dimensional regression or classification frameworks, variable selection is a difficult task, that becomes even more…

统计方法学 · 统计学 2016-04-19 Baptiste Gregorutti , Bertrand Michel , Philippe Saint-Pierre

We propose a modification that corrects for split-improvement variable importance measures in Random Forests and other tree-based methods. These methods have been shown to be biased towards increasing the importance of features with more…

机器学习 · 统计学 2020-03-25 Zhengze Zhou , Giles Hooker

Reliable estimation of feature contributions in machine learning models is essential for trust, transparency and regulatory compliance, especially when models are proprietary or otherwise operate as black boxes. While permutation-based…

机器学习 · 统计学 2025-12-24 Albert Dorador

A common problem in machine learning is determining if a variable significantly contributes to a model's prediction performance. This problem is aggravated for datasets, such as gene expression datasets, that suffer the worst case of…

统计方法学 · 统计学 2023-10-13 Yue Wu , Ted Spaide , Kenji Nakamichi , Russell Van Gelder , Aaron Lee

Along with accurate prediction, understanding the contribution of each feature to the making of the prediction, i.e., the importance of the feature, is a desirable and arguably necessary component of a machine learning model. For a complex…

机器学习 · 计算机科学 2025-07-11 Aaron Foote , Danny Krizanc

The default variable-importance measure in random Forests, Gini importance, has been shown to suffer from the bias of the underlying Gini-gain splitting criterion. While the alternative permutation importance is generally accepted as a…

机器学习 · 统计学 2020-05-18 Markus Loecher

We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally…

机器学习 · 统计学 2009-09-29 Hemant Ishwaran

Random Forest is a machine learning method that offers many advantages, including the ability to easily measure variable importance. Class balancing technique is a well-known solution to deal with class imbalance problem. However, it has…

机器学习 · 统计学 2023-12-19 Yunbi Nam , Sunwoo Han

Random Forests have become a widely used tool in machine learning since their introduction in 2001, known for their strong performance in classification and regression tasks. One key feature of Random Forests is the Random Forest…

统计理论 · 数学 2025-12-18 Nico Föge , Lena Schmid , Marc Ditzhaus , Markus Pauly

Variable importance assessment has become a crucial step in machine-learning applications when using complex learners, such as deep neural networks, on large-scale data. Removal-based importance assessment is currently the reference…

机器学习 · 计算机科学 2023-10-27 Ahmad Chamma , Denis A. Engemann , Bertrand Thirion

Hypothesis testing of random forest (RF) variable importance measures (VIMP) remains the subject of ongoing research. Among recent developments, heuristic approaches to parametric testing have been proposed whose distributional assumptions…

统计方法学 · 统计学 2023-07-20 Alexander Hapfelmeier , Roman Hornung , Bernhard Haller

While achieving high prediction accuracy is a fundamental goal in machine learning, an equally important task is finding a small number of features with high explanatory power. One popular selection technique is permutation importance,…

机器学习 · 统计学 2024-10-02 Min Lu , Hemant Ishwaran

Random forests are one of the most popular machine learning methods due to their accuracy and variable importance assessment. However, random forests only provide variable importance in a global sense. There is an increasing need for such…

统计方法学 · 统计学 2021-03-25 Joshua Daniel Loyal , Ruoqing Zhu , Yifan Cui , Xin Zhang

Variable selection is an important statistical problem. This problem becomes more challenging when the candidate predictors are of mixed type (e.g. continuous and binary) and impact the response variable in nonlinear and/or non-additive…

统计方法学 · 统计学 2021-12-30 Chuji Luo , Michael J. Daniels

Random Forests are renowned for their predictive accuracy, but valid inference, particularly about permutation-based feature importances, remains challenging. Existing methods, such as the confidence intervals (CIs) from Ishwaran et al.…

统计方法学 · 统计学 2025-07-21 Nico Föge , Markus Pauly

Quantile Regression Forests (QRF) are widely used for non-parametric conditional quantile estimation, yet statistical inference for variable importance measures remains challenging due to the non-smoothness of the loss function and the…

机器学习 · 统计学 2025-12-01 Tomoshige Nakamura , Hiroshi Shiraishi

In order to trust the predictions of a machine learning algorithm, it is necessary to understand the factors that contribute to those predictions. In the case of probabilistic and uncertainty-aware models, it is necessary to understand not…

机器学习 · 统计学 2024-08-19 Danny Wood , Theodore Papamarkou , Matt Benatan , Richard Allmendinger

This paper introduces and develops a novel variable importance score function in the context of ensemble learning and demonstrates its appeal both theoretically and empirically. Our proposed score function is simple and more straightforward…

机器学习 · 统计学 2015-01-27 Ernest Fokoué

New inference methods for the multivariate coefficient of variation and its reciprocal, the standardized mean, are presented. While there are various testing procedures for both parameters in the univariate case, it is less known how to do…

统计方法学 · 统计学 2020-03-31 Marc Ditzhaus , Łukas Smaga
‹ 上一页 1 2 3 10 下一页 ›