English

A Novel Taxonomy for Navigating and Classifying Synthetic Data in Healthcare Applications

Computers and Society 2024-09-04 v1

Abstract

Data-driven technologies have improved the efficiency, reliability and effectiveness of healthcare services, but come with an increasing demand for data, which is challenging due to privacy-related constraints on sharing data in healthcare contexts. Synthetic data has recently gained popularity as potential solution, but in the flurry of current research it can be hard to oversee its potential. This paper proposes a novel taxonomy of synthetic data in healthcare to navigate the landscape in terms of three main varieties. Data Proportion comprises different ratios of synthetic data in a dataset and associated pros and cons. Data Modality refers to the different data formats amenable to synthesis and format-specific challenges. Data Transformation concerns improving specific aspects of a dataset like its utility or privacy with synthetic data. Our taxonomy aims to help researchers in the healthcare domain interested in synthetic data to grasp what types of datasets, data modalities, and transformations are possible with synthetic data, and where the challenges and overlaps between the varieties lie.

Keywords

Cite

@article{arxiv.2409.00701,
  title  = {A Novel Taxonomy for Navigating and Classifying Synthetic Data in Healthcare Applications},
  author = {Bram van Dijk and Saif ul Islam and Jim Achterberg and Hafiz Muhammad Waseem and Parisis Gallos and Gregory Epiphaniou and Carsten Maple and Marcel Haas and Marco Spruit},
  journal= {arXiv preprint arXiv:2409.00701},
  year   = {2024}
}

Comments

Accepted at the 23rd EFMI Special Topic Conference, Romania, November 2024