中文
相关论文

相关论文: Synthetic Tabular Data Generation: A Comparative S…

200 篇论文

Deep generative models and synthetic medical data have shown significant promise in addressing key challenges in healthcare, such as privacy concerns, data bias, and the scarcity of realistic datasets. While research in this area has grown…

机器学习 · 计算机科学 2025-02-05 Krishan Agyakari Raja Babu , Supriti Mulay , Om Prabhu , Mohanasankar Sivaprakasam

Institutions collect massive learning traces but they may not disclose it for privacy issues. Synthetic data generation opens new opportunities for research in education. In this paper we present a generative model for educational data that…

计算机与社会 · 计算机科学 2022-07-09 Jill-Jênn Vie , Tomas Rigaux , Sein Minn

Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong…

机器学习 · 计算机科学 2026-02-09 Maj-Annika Tammisto , Faiz Ali Shah , Daniel Rodriguez , Dietmar Pfahl

AI-based data synthesis has seen rapid progress over the last several years, and is increasingly recognized for its promise to enable privacy-respecting high-fidelity data sharing. However, adequately evaluating the quality of generated…

机器学习 · 统计学 2021-04-02 Michael Platzer , Thomas Reutterer

Tabular synthesis models remain ineffective at capturing complex dependencies, and the quality of synthetic data is still insufficient for comprehensive downstream tasks, such as prediction under distribution shifts, automated…

机器学习 · 计算机科学 2024-07-08 Ruibo Tu , Zineb Senane , Lele Cao , Cheng Zhang , Hedvig Kjellström , Gustav Eje Henter

Synthetic tabular data generation addresses data scarcity and privacy constraints in a variety of domains. Tabular Prior-Data Fitted Network (TabPFN), a recent foundation model for tabular data, has been shown capable of generating…

机器学习 · 计算机科学 2026-03-12 Davide Tugnoli , Andrea De Lorenzo , Marco Virgolin , Giovanni Cinà

We present a novel approach for differentially private data synthesis of protected tabular datasets, a relevant task in highly sensitive domains such as healthcare and government. Current state-of-the-art methods predominantly use…

机器学习 · 计算机科学 2024-07-30 Konstantin Donhauser , Javier Abad , Neha Hulkund , Fanny Yang

Synthetic Data is increasingly important in financial applications. In addition to the benefits it provides, such as improved financial modeling and better testing procedures, it poses privacy risks as well. Such data may arise from client…

密码学与安全 · 计算机科学 2024-03-25 Tucker Balch , Vamsi K. Potluru , Deepak Paramanand , Manuela Veloso

Synthetically generated data can improve privacy, fairness, and data accessibility; however, it can be challenging in specialized scenarios such as survival analysis. One key challenge in this setting is censoring, i.e., the timing of an…

机器学习 · 统计学 2025-08-07 Mohd Ashhad , Ricardo Henao

Exploiting the recent advancements in artificial intelligence, showcased by ChatGPT and DALL-E, in real-world applications necessitates vast, domain-specific, and publicly accessible datasets. Unfortunately, the scarcity of such datasets…

机器学习 · 计算机科学 2023-05-17 Cyril Picard , Jürg Schiffmann , Faez Ahmed

Probabilistic relational models provide a well-established formalism to combine first-order logic and probabilistic models, thereby allowing to represent relationships between objects in a relational domain. At the same time, the field of…

人工智能 · 计算机科学 2024-10-03 Malte Luttermann , Ralf Möller , Mattis Hartwig

Synthetic tabular data is becoming a necessity as concerns about data privacy intensify in the world. Tabular data can be useful for testing various systems, simulating real data, analyzing the data itself or building predictive models.…

机器学习 · 计算机科学 2024-07-19 Volodymyr Shulakov

Sharing sensitive data is vital in enabling many modern data analysis and machine learning tasks. However, current methods for data release are insufficiently accurate or granular to provide meaningful utility, and they carry a high risk of…

数据库 · 计算机科学 2021-08-25 Teddy Cunningham , Graham Cormode , Hakan Ferhatosmanoglu

Synthetic data is often perceived as a silver-bullet solution to data anonymization and privacy-preserving data publishing. Drawn from generative models like diffusion models, synthetic data is expected to preserve the statistical…

Generative data augmentation (GDA) has emerged as a promising technique to alleviate data scarcity in machine learning applications. This thesis presents a comprehensive survey and unified framework of the GDA landscape. We first provide an…

机器学习 · 计算机科学 2024-04-23 Yunhao Chen , Zihui Yan , Yunjie Zhu

In the rapidly evolving field of artificial intelligence, the creation and utilization of synthetic datasets have become increasingly significant. This report delves into the multifaceted aspects of synthetic data, particularly emphasizing…

机器学习 · 计算机科学 2024-01-04 Shuang Hao , Wenfeng Han , Tao Jiang , Yiping Li , Haonan Wu , Chunlin Zhong , Zhangjun Zhou , He Tang

We propose a method to generate statistically representative synthetic data from a given dataset. The main goal of our method is for the created data set to mimic the inter--feature correlations present in the original data, while also…

机器学习 · 计算机科学 2025-06-25 Nicklas Jävergård , Rainey Lyons , Adrian Muntean , Jonas Forsman

We provide new algorithms for two tasks relating to heterogeneous tabular datasets: clustering, and synthetic data generation. Tabular datasets typically consist of heterogeneous data types (numerical, ordinal, categorical) in columns, but…

机器学习 · 计算机科学 2024-04-22 Chandrani Kumari , Rahul Siddharthan

Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes. In…

机器学习 · 计算机科学 2026-05-28 Junfeng Nie , Alvin Jin , Xiaohui Chen

Artificial intelligence and data access are already mainstream. One of the main challenges when designing an artificial intelligence or disclosing content from a database is preserving the privacy of individuals who participate in the…

密码学与安全 · 计算机科学 2023-12-13 Clément Pierquin , Bastien Zimmermann , Matthieu Boussard