English
Related papers

Related papers: A Framework for Generating Realistic Synthetic Tab…

200 papers

Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. Mainstream research paradigm necessitates real-world person images with manual textual annotations for training models, posing privacy concerns and annotation…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Min Cao , Yuxin Lu , Ziyin Zeng , Dong Yi , Jinqiao Wang , Mang Ye

Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have…

Artificial Intelligence · Computer Science 2024-02-01 Katherine A. Keith , Sergey Feldman , David Jurgens , Jonathan Bragg , Rohit Bhattacharya

Tabular data is one of the most widely used data formats across various domains such as bioinformatics, healthcare, and marketing. As artificial intelligence moves towards a data-centric perspective, improving data quality is essential for…

Generating clinical synthetic text represents an effective solution for common clinical NLP issues like sparsity and privacy. This paper aims to conduct a systematic review on generating synthetic medical free-text by formulating…

Computation and Language · Computer Science 2025-07-25 Basel Alshaikhdeeb , Ahmed Abdelmonem Hemedan , Soumyabrata Ghosh , Irina Balaur , Venkata Satagopam

Recent advances in deep generative models have greatly expanded the potential to create realistic synthetic health datasets. These synthetic datasets aim to preserve the characteristics, patterns, and overall scientific conclusions derived…

Machine Learning · Computer Science 2024-07-04 Jennifer A Bartell , Sander Boisen Valentin , Anders Krogh , Henning Langberg , Martin Bøgsted

AI requires extensive datasets, while medical data is subject to high data protection. Anonymization is essential, but poses a challenge for some regions, such as the head, as identifying structures overlap with regions of clinical…

Synthetic data is being used lately for training deep neural networks in computer vision applications such as object detection, object segmentation and 6D object pose estimation. Domain randomization hereby plays an important role in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Parth Rawal , Mrunal Sompura , Wolfgang Hintze

Generating datasets that "look like" given real ones is an interesting tasks for healthcare applications of ML and many other fields of science and engineering. In this paper we propose a new method of general application to binary datasets…

Machine Learning · Statistics 2018-07-05 Laura Aviñó , Matteo Ruffini , Ricard Gavaldà

Randomised controlled trials (RCTs) are regarded as the gold standard for estimating causal treatment effects on health outcomes. However, RCTs are not always feasible, because of time, budget or ethical constraints. Observational data such…

Methodology · Statistics 2024-02-20 Li Su , Roonak Rezvani , Shaun R. Seaman , Colin Starr , Isaac Gravestock

There is no consensus in the field of synthetic data on concise metrics for quality evaluations or benchmarks on large health datasets, such as historical epidemiological data. This study presents an evaluation of seven recent models from…

Machine Learning · Computer Science 2026-04-20 Jean-Baptiste Escudié , Benjamin Barnes , Stefan Meisegeier , Klaus Kraywinkel , Fabian Prasser , Nils Körber

Data sharing is a prerequisite for collaborative innovation, enabling organizations to leverage diverse datasets for deeper insights. In real-world applications like FinTech and Smart Manufacturing, transactional data, often in tabular…

Cryptography and Security · Computer Science 2024-11-07 Mengmeng Yang , Chi-Hung Chi , Kwok-Yan Lam , Jie Feng , Taolin Guo , Wei Ni

Observational studies provide the only evidence on the effectiveness of interventions when randomized controlled trials (RCTs) are impractical due to cost, ethical concerns, or time constraints. While many methodologies aim to draw causal…

Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes. In…

Machine Learning · Computer Science 2026-05-28 Junfeng Nie , Alvin Jin , Xiaohui Chen

Generative Adversarial Networks (GANs) have shown remarkable success as a framework for training models to produce realistic-looking data. In this work, we propose a Recurrent GAN (RGAN) and Recurrent Conditional GAN (RCGAN) to produce…

Machine Learning · Statistics 2017-12-05 Cristóbal Esteban , Stephanie L. Hyland , Gunnar Rätsch

Privacy-preserving synthetic data offers a promising solution to harness segregated data in high-stakes domains where information is compartmentalized for regulatory, privacy, or institutional reasons. This survey provides a comprehensive…

Cryptography and Security · Computer Science 2025-03-28 Viktor Schlegel , Anil A Bharath , Zilong Zhao , Kevin Yee

The widespread adoption of electronic health records and digital healthcare data has created a demand for data-driven insights to enhance patient outcomes, diagnostics, and treatments. However, using real patient data presents privacy and…

Machine Learning · Computer Science 2023-11-15 Aryan Jadon , Shashank Kumar

Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and…

Synthetic tabular data generation is increasingly essential in data management, supporting downstream applications when real-world and high-quality tabular data is insufficient. Existing tabular generation approaches, such as generative…

Machine Learning · Computer Science 2025-09-15 Mingxuan Jiang , Yongxin Wang , Ziyue Dai , Yicun Liu , Hongyi Nie , Sen Liu , Hongfeng Chai

Clinical data usually cannot be freely distributed due to their highly confidential nature and this hampers the development of machine learning in the healthcare domain. One way to mitigate this problem is by generating realistic synthetic…

We provide new algorithms for two tasks relating to heterogeneous tabular datasets: clustering, and synthetic data generation. Tabular datasets typically consist of heterogeneous data types (numerical, ordinal, categorical) in columns, but…

Machine Learning · Computer Science 2024-04-22 Chandrani Kumari , Rahul Siddharthan
‹ Prev 1 3 4 5 6 7 10 Next ›