English
Related papers

Related papers: Memisis: Orchestrating and Evaluating Synthetic Da…

200 papers

This paper presents SYMBIOSIS, an AI-powered framework and platform designed to make Systems Thinking accessible for addressing societal challenges and unlock paths for leveraging systems thinking frameworks to improve AI systems. The…

Computers and Society · Computer Science 2025-03-11 Sameer Sethi , Donald Martin , Emmanuel Klu

Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, existing pipelines remain fragmented and…

Ensuring safe adoption of AI tools in healthcare hinges on access to sufficient data for training, testing and validation. In response to privacy concerns and regulatory requirements, using synthetic data has been suggested. Synthetic data…

Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge…

Machine Learning · Computer Science 2026-02-19 Nitish Nagesh , Salar Shakibhamedan , Mahdi Bagheri , Ziyu Wang , Nima TaheriNejad , Axel Jantsch , Amir M. Rahmani

Synthetic tabular data generation becomes crucial when real data is limited, expensive to collect, or simply cannot be used due to privacy concerns. However, producing good quality synthetic data is challenging. Several probabilistic,…

Machine Learning · Computer Science 2024-06-11 Vikram S Chundawat , Ayush K Tarun , Murari Mandal , Mukund Lahoti , Pratik Narang

Faced with the challenges of patient confidentiality and scientific reproducibility, research on machine learning for health is turning towards the conception of synthetic medical databases. This article presents a brief overview of…

Synthetic data has emerged as a powerful resource in life sciences, offering solutions for data scarcity, privacy protection and accessibility constraints. By creating artificial datasets that mirror the characteristics of real data, allows…

Synthcity is an open-source software package for innovative use cases of synthetic data in ML fairness, privacy and augmentation across diverse tabular data modalities, including static data, regular and irregular time series, data with…

Machine Learning · Computer Science 2023-01-19 Zhaozhi Qian , Bogdan-Constantin Cebere , Mihaela van der Schaar

With the advent of generative modeling techniques, synthetic data and its use has penetrated across various domains from unstructured data such as image, text to structured dataset modeling healthcare outcome, risk decisioning in financial…

Machine Learning · Computer Science 2021-05-11 Aman Gupta , Deepak Bhatt , Anubha Pandey

Access to electronic health records (EHRs) for digital health research is often limited by privacy regulations and institutional barriers. Synthetic EHRs have been proposed as a way to enable safe and sovereign data sharing; however,…

Machine Learning · Computer Science 2026-03-10 Guanglin Zhou , Armin Catic , Motahare Shabestari , Matthew Young , Chaiquan Li , Katrina Poppe , Sebastiano Barbieri

This paper explores the strategic use of modern synthetic data generation and advanced data perturbation techniques to enhance security, maintain analytical utility, and improve operational efficiency when managing large datasets, with a…

Cryptography and Security · Computer Science 2025-04-29 Anantha Sharma , Swetha Devabhaktuni , Eklove Mohan

Synthetic Data Generation (SDG) based on Artificial Intelligence (AI) can transform the way clinical medicine is delivered by overcoming privacy barriers that currently render clinical data sharing difficult. This is the key to accelerating…

The rapid advancements in generative AI and large language models (LLMs) have opened up new avenues for producing synthetic data, particularly in the realm of structured tabular formats, such as product reviews. Despite the potential…

Machine Learning · Computer Science 2025-07-25 Yefeng Yuan , Yuhong Liu , Liang Cheng

Data for good implies unfettered access to data. But data owners must be conservative about how, when, and why they share data or risk violating the trust of the people they aim to help, losing their funding, or breaking the law. Data…

Computers and Society · Computer Science 2017-10-25 Bill Howe , Julia Stoyanovich , Haoyue Ping , Bernease Herman , Matt Gee

Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate unbiased, privacy-preserving data while maintaining…

Existing pose estimation models perform poorly on wheelchair users due to a lack of representation in training data. We present a data synthesis pipeline to address this disparity in data collection and subsequently improve pose estimation…

Human-Computer Interaction · Computer Science 2026-01-23 William Huang , Sam Ghahremani , Siyou Pei , Yang Zhang

Introduction: The amount of data generated by original research is growing exponentially. Publicly releasing them is recommended to comply with the Open Science principles. However, data collected from human participants cannot be released…

Machine Learning · Statistics 2023-10-11 Rémy Chapelle , Bruno Falissard

We provide new algorithms for two tasks relating to heterogeneous tabular datasets: clustering, and synthetic data generation. Tabular datasets typically consist of heterogeneous data types (numerical, ordinal, categorical) in columns, but…

Machine Learning · Computer Science 2024-04-22 Chandrani Kumari , Rahul Siddharthan

This research delves into the construction and utilization of synthetic datasets, specifically within the telematics sphere, leveraging OpenAI's powerful language model, ChatGPT. Synthetic datasets present an effective solution to…

Computers and Society · Computer Science 2023-06-27 Ryan Lingo

Synthetic datasets are important for evaluating and testing machine learning models. When evaluating real-life recommender systems, high-dimensional categorical (and sparse) datasets are often considered. Unfortunately, there are not many…

Information Retrieval · Computer Science 2024-12-11 Miha Malenšek , Blaž Škrlj , Blaž Mramor , Jure Demšar