面向官方统计的机器学习:统计学宣言
机器学习
2024-09-09 v1 机器学习
统计方法学
摘要
官方统计生产应用机器学习必须具有统计学严谨性,因为它既带来机遇,也面临挑战。尽管机器学习在近年来取得了快速技术进步,但其应用方法尚不具备产生高质量统计结果所必需的鲁棒性。为了消除机器学习模型中所有误差来源,本文提出了类似于调查方法学中总体调查误差模型的总体机器学习误差 (TMLE) 框架。作为确保机器学习模型既在内部有效又在外部有效的手段,TMLE 模型解决了代表性问题和测量误差等问题。本文还提供了多个案例研究,说明在官方统计中更严格地应用机器学习的重要性。
引用
@article{arxiv.2409.04365,
title = {Leveraging Machine Learning for Official Statistics: A Statistical Manifesto},
author = {Marco Puts and David Salgado and Piet Daas},
journal= {arXiv preprint arXiv:2409.04365},
year = {2024}
}
备注
29 pages, 4 figures, 1 table. To appear in the proceedings of the conference on Foundations and Advances of Machine Learning in Official Statistics, which was held in Wiesbaden, from 3rd to 5th April, 2024