中文

用于多发性硬化症生物标志物发现的机器学习流水线:可解释人工智能与传统统计方法的比较

机器学习 2025-09-29 v1 人工智能

摘要

我们提出了一种用于多发性硬化症(Multiple Sclerosis, MS)生物标志物发现的机器学习流水线,整合了来自外周血单核细胞(Peripheral Blood Mononuclear Cells, PBMC)的八个公开微阵列数据集。在稳健的预处理之后,我们训练了一个通过贝叶斯搜索优化的 XGBoost 分类器。我们使用 SHapley Additive exPlanations (SHAP) 来识别模型预测的关键特征,从而指示可能的生物标志物。我们将这些特征与通过经典差异表达分析(Differential Expression Analysis, DEA)识别出的基因进行了比较。我们的比较揭示了 SHAP 和 DEA 之间既有重叠也有独特的生物标志物,表明它们具有互补的优势。富集分析证实了 SHAP 所选基因的生物学相关性,将其与鞘脂信号传导、Th1/Th2/Th17 细胞分化以及 Epstein-Barr 病毒感染等已知与 MS 相关的通路联系起来。本研究强调了将可解释人工智能(xAI)与传统统计方法相结合的价值,以更深入地了解疾病机制。

关键词

引用

@article{arxiv.2509.22484,
  title  = {A Machine Learning Pipeline for Multiple Sclerosis Biomarker Discovery: Comparing explainable AI and Traditional Statistical Approaches},
  author = {Samuele Punzo and Silvia Giulia Galfrè and Francesco Massafra and Alessandro Maglione and Corrado Priami and Alina Sîrbu},
  journal= {arXiv preprint arXiv:2509.22484},
  year   = {2025}
}

备注

Short paper presented at the 20th conference on Computational Intelligence methods for Bioinformatics and Biostatistics (CIBB2025)