English
Related papers

Related papers: Transitioning from Real to Synthetic data: Quantif…

200 papers

Computational text classification is a challenging task, especially for multi-dimensional social constructs. Recently, there has been increasing discussion that synthetic training data could enhance classification by offering examples of…

Computation and Language · Computer Science 2024-12-11 Lukas Birkenmaier , Matthias Roth , Indira Sen

With promising empirical performance across a wide range of applications, synthetic data augmentation appears a viable solution to data scarcity and the demands of increasingly data-intensive models. Its effectiveness lies in expanding the…

Machine Learning · Computer Science 2026-02-02 Zixuan Wu , So Won Jeong , Yating Liu , Yeo Jin Jung , Claire Donnat

The current literature regarding generation of complex, realistic synthetic tabular data, particularly for randomized controlled trials (RCTs), often ignores missing data. However, missing data are common in RCT data and often are not…

Other Statistics · Statistics 2025-12-02 Niki Z. Petrakos , Erica E. M. Moodie , Nicolas Savy

Dataset bias is a significant challenge in machine learning, where specific attributes, such as texture or color of the images are unintentionally learned resulting in detrimental performance. To address this, previous efforts have focused…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Donggeun Ko , Sangwoo Jo , Dongjun Lee , Namjun Park , Jaekwang Kim

The availability of genomic data is essential to progress in biomedical research, personalized medicine, etc. However, its extreme sensitivity makes it problematic, if not outright impossible, to publish or share it. As a result, several…

Genomics · Quantitative Biology 2022-01-19 Bristena Oprisanu , Georgi Ganev , Emiliano De Cristofaro

Generative models unfairly penalize data belonging to minority classes, suffer from model autophagy disorder (MADness), and learn biased estimates of the underlying distribution parameters. Our theoretical and empirical results show that…

Machine Learning · Computer Science 2024-10-07 Paul Mayer , Lorenzo Luzi , Ali Siahkoohi , Don H. Johnson , Richard G. Baraniuk

Synthetic data generation is important to training and evaluating neural models for question answering over knowledge graphs. The quality of the data and the partitioning of the datasets into training, validation and test splits impact the…

Information Retrieval · Computer Science 2020-09-11 Trond Linjordet , Krisztian Balog

Among various aspects of ensuring the responsible design of AI tools for healthcare applications, addressing fairness concerns has been a key focus area. Specifically, given the wide spread of electronic health record (EHR) data and their…

Machine Learning · Computer Science 2025-06-30 Mirza Farhan Bin Tarek , Raphael Poulain , Rahmatollah Beheshti

Synthetic data generation is gaining traction as a privacy enhancing technology (PET). When properly generated, synthetic data preserve the analytic utility of real data while avoiding the retention of information that would allow the…

It has been demonstrated that the amount of data is crucial in data-driven machine learning methods. Data is always valuable, but in some tasks, it is almost like gold. This occurs in engineering areas where data is scarce or very expensive…

Artificial Intelligence · Computer Science 2023-12-12 David Solis-Martin , Juan Galan-Paez , Joaquin Borrego-Diaz

Big data analysis poses the dual problem of privacy preservation and utility, i.e., how accurate data analyses remain after transforming original data in order to protect the privacy of the individuals that the data is about - and whether…

Machine Learning · Computer Science 2022-11-29 Md Sakib Nizam Khan , Niklas Reje , Sonja Buchegger

Natural Language Processing (NLP) has undergone transformative changes with the advent of deep learning methodologies. One challenge persistently confronting researchers is the scarcity of high-quality, annotated datasets that drive these…

Computation and Language · Computer Science 2023-10-13 Sia Gholami , Marwan Omar

Training medical AI algorithms requires large volumes of accurately labeled datasets, which are difficult to obtain in the real world. Synthetic images generated from deep generative models can help alleviate the data scarcity problem, but…

Image and Video Processing · Electrical Eng. & Systems 2023-06-16 Xiaodan Xing , Yang Nan , Federico Felder , Simon Walsh , Guang Yang

Although many fairness criteria have been proposed to ensure that machine learning algorithms do not exhibit or amplify our existing social biases, these algorithms are trained on datasets that can themselves be statistically biased. In…

Machine Learning · Computer Science 2023-05-04 Yiqiao Liao , Parinaz Naghizadeh

Exploiting the recent advancements in artificial intelligence, showcased by ChatGPT and DALL-E, in real-world applications necessitates vast, domain-specific, and publicly accessible datasets. Unfortunately, the scarcity of such datasets…

Machine Learning · Computer Science 2023-05-17 Cyril Picard , Jürg Schiffmann , Faez Ahmed

Data augmentations are useful in closing the sim-to-real domain gap when training on synthetic data. This is because they widen the training data distribution, thus encouraging the model to generalize better to other domains. Many image…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Bram Vanherle , Nick Michiels , Frank Van Reeth

In this paper, we address a key scientific problem in machine learning: Given a training set for an image classification task, can we train a generative model on this dataset to enhance the classification performance? (i.e., closed-set…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Haowen Wang , Guowei Zhang , Xiang Zhang , Zeyuan Chen , Haiyang Xu , Dou Hoon Kwark , Zhuowen Tu

Synthetic images generated from deep generative models have the potential to address data scarcity and data privacy issues. The selection of synthesis models is mostly based on image quality measurements, and most researchers favor…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Xiaodan Xing , Federico Felder , Yang Nan , Giorgos Papanastasiou , Walsh Simon , Guang Yang

Machine learning has the potential to assist many communities in using the large datasets that are becoming more and more available. Unfortunately, much of that potential is not being realized because it would require sharing data in a way…

Machine Learning · Computer Science 2018-07-02 James Jordon , Jinsung Yoon , Mihaela van der Schaar

We explore the privacy-utility tradeoff of synthetic data generation schemes on tabular financial datasets, a domain characterized by high regulatory risk and severe class imbalance. We consider representative tabular data generators,…

Machine Learning · Computer Science 2026-02-11 Michael Zuo , Inwon Kang , Stacy Patterson , Oshani Seneviratne
‹ Prev 1 8 9 10 Next ›