中文

PrInDT 中的重复欠采样(RePrInDT):欠采样与预测的变异性及集成中预测变量的排序

应用统计 2021-08-12 v1

摘要

在本文中,我们将我们的 PrInDT 方法(Weihs & Buschfeld 2021a)扩展至较小类和较大类不同百分比(psmall 和 plarge)的欠采样、预测变量的分层、预测阈值的变动,以及集成中变量重要性的度量。将这些方法应用于一个语言学示例表明:1. 在欠采样中,仔细选择百分比 plarge 和 psmall 对于构建具有高平衡准确率的模型很重要;2. 预测变量的分层并未大幅提升平衡准确率;3. 降低较小类的预测阈值被证明是欠采样的一种替代方法,因为它增加了较小类被选中的可能性。最后,我们引入一种对预测变量重要性进行排序的方法,可对结果进行直观解释。

关键词

引用

@article{arxiv.2108.05129,
  title  = {Repeated undersampling in PrInDT (RePrInDT): Variation in undersampling and prediction, and ranking of predictors in ensembles},
  author = {Claus Weihs and Sarah Buschfeld},
  journal= {arXiv preprint arXiv:2108.05129},
  year   = {2021}
}