弥合泛化鸿沟:在混杂生物数据上训练鲁棒模型
机器学习
2018-12-13 v1 机器学习
摘要
由于样本采集与处理中的混杂变量,在生物数据上进行统计学习可能具有挑战性。若模型未经过彻底验证,混杂因素会导致模型泛化能力差,并造成不准确的预测性能指标。在本文中,我们提出控制混杂因素并进一步提升预测性能的方法。我们引入混杂因子归一化中的正交基构造(ONION, OrthoNormal basis construction In cOnfounding factor Normalization)以消除混杂协变量,并使用域对抗神经网络(DANN, Domain-Adversarial Neural Network)来惩罚对混杂信息进行编码的模型。我们将所提方法应用于模拟与实证患者数据,并展示了泛化能力的显著提升。
引用
@article{arxiv.1812.04778,
title = {Bridging the Generalization Gap: Training Robust Models on Confounded Biological Data},
author = {Tzu-Yu Liu and Ajay Kannan and Adam Drake and Marvin Bertin and Nathan Wan},
journal= {arXiv preprint arXiv:1812.04778},
year = {2018}
}