English

Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data

Machine Learning 2025-07-08 v1 Artificial Intelligence Methodology Machine Learning

Abstract

Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark.

Keywords

Cite

@article{arxiv.2507.03971,
  title  = {Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data},
  author = {Anurag Garg and Muhammad Ali and Noah Hollmann and Lennart Purucker and Samuel Müller and Frank Hutter},
  journal= {arXiv preprint arXiv:2507.03971},
  year   = {2025}
}
R2 v1 2026-07-01T03:47:34.879Z