English

Leveraging Machine Learning for Official Statistics: A Statistical Manifesto

Machine Learning 2024-09-09 v1 Machine Learning Methodology

Abstract

It is important for official statistics production to apply ML with statistical rigor, as it presents both opportunities and challenges. Although machine learning has enjoyed rapid technological advances in recent years, its application does not possess the methodological robustness necessary to produce high quality statistical results. In order to account for all sources of error in machine learning models, the Total Machine Learning Error (TMLE) is presented as a framework analogous to the Total Survey Error Model used in survey methodology. As a means of ensuring that ML models are both internally valid as well as externally valid, the TMLE model addresses issues such as representativeness and measurement errors. There are several case studies presented, illustrating the importance of applying more rigor to the application of machine learning in official statistics.

Keywords

Cite

@article{arxiv.2409.04365,
  title  = {Leveraging Machine Learning for Official Statistics: A Statistical Manifesto},
  author = {Marco Puts and David Salgado and Piet Daas},
  journal= {arXiv preprint arXiv:2409.04365},
  year   = {2024}
}

Comments

29 pages, 4 figures, 1 table. To appear in the proceedings of the conference on Foundations and Advances of Machine Learning in Official Statistics, which was held in Wiesbaden, from 3rd to 5th April, 2024

R2 v1 2026-06-28T18:36:38.063Z