English

On the Usefulness of Synthetic Tabular Data Generation

Machine Learning 2023-06-28 v1

Abstract

Despite recent advances in synthetic data generation, the scientific community still lacks a unified consensus on its usefulness. It is commonly believed that synthetic data can be used for both data exchange and boosting machine learning (ML) training. Privacy-preserving synthetic data generation can accelerate data exchange for downstream tasks, but there is not enough evidence to show how or why synthetic data can boost ML training. In this study, we benchmarked ML performance using synthetic tabular data for four use cases: data sharing, data augmentation, class balancing, and data summarization. We observed marginal improvements for the balancing use case on some datasets. However, we conclude that there is not enough evidence to claim that synthetic tabular data is useful for ML training.

Keywords

Cite

@article{arxiv.2306.15636,
  title  = {On the Usefulness of Synthetic Tabular Data Generation},
  author = {Dionysis Manousakas and Sergül Aydöre},
  journal= {arXiv preprint arXiv:2306.15636},
  year   = {2023}
}

Comments

Data-centric Machine Learning Research (DMLR) Workshop at the 40th International Conference on Machine Learning (ICML)