中文

预测建模中混杂偏倚的统计量化

机器学习 2025-05-30 v1 定量方法 机器学习

摘要

缺乏针对混杂偏倚的非参数统计检验,严重阻碍了众多研究领域中稳健、有效且可泛化预测模型的发展。在此,我提出偏混杂检验与全混杂检验,对于给定的混杂变量,分别探查无混杂模型与完全混杂模型的原假设。这些检验对 I 型错误提供严格控制并具有高统计功效,即便对于机器学习中常见的非正态与非线性依赖预测亦如此。将所提检验应用于基于 Human Connectome Project 与 Autism Brain Imaging Data Exchange 数据集的功能脑连接数据训练的模型,揭示了先前未报道或被认为难以用最先进混杂缓解方法纠正的混杂因素。这些检验在 mlconfound 包中实现 (https://mlconfound.readthedocs.io),可辅助评估和提升预测模型的泛化能力与神经生物学有效性,从而推动临床有用机器学习生物标志物的发展。

关键词

引用

@article{arxiv.2111.00814,
  title  = {Statistical quantification of confounding bias in predictive modelling},
  author = {Tamas Spisak},
  journal= {arXiv preprint arXiv:2111.00814},
  year   = {2025}
}

备注

20 pages, 7 figures. The manuscript is associated with the the python package `mlconfound`: https://mlconfound.readthedocs.io See manuscript repository, including fully reproducible analysis code, here: https://github.com/pni-lab/mlconfound-manuscript