Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.
@article{arxiv.2511.04000,
title = {Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations},
author = {Kyaw Hpone Myint and Zhe Wu and Alexandre G. R. Day and Giri Iyengar},
journal= {arXiv preprint arXiv:2511.04000},
year = {2025}
}
Comments
9 pages, 3 figures, Neurips 2025 GenAI in Finance Workshop