English
Related papers

Related papers: Generative Synthesis of Insurance Datasets

200 papers

When working with real-world insurance data, practitioners often encounter challenges during the data preparation stage that can undermine the statistical validity and reliability of downstream modeling. This study illustrates that…

Machine Learning · Statistics 2026-03-20 Jiayi Guo , Panyi Dong , Zhiyu Quan

Generative AI workflows heavily rely on data-centric tasks - such as filtering samples by annotation fields, vector distances, or scores produced by custom classifiers. At the same time, computer vision datasets are quickly approaching…

Artificial Intelligence · Computer Science 2023-09-22 Daniel Kharitonov , Ryan Turner

Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate unbiased, privacy-preserving data while maintaining…

Generative models for financial time series often create data that look realistic and even reproduce stylized facts such as fat tails or volatility clustering. However, these apparent successes break down under trading backtests: models…

Statistical Finance · Quantitative Finance 2026-01-21 Fan Zhang , Jiabin Luo , Zheng Zhang , Shuanghong Huang , Zhipeng Liu , Yu Chen

The emergence of generative AI models has dramatically expanded the availability and use of synthetic data across scientific, industrial, and policy domains. While these developments open new possibilities for data analysis, they also raise…

Machine Learning · Statistics 2026-03-06 Ahmad Abdel-Azim , Ruoyu Wang , Xihong Lin

A novel strategy for generating datasets is developed within the context of drag prediction for automotive geometries using neural networks. A primary challenge in this space is constructing a training databse of sufficient size and…

Machine Learning · Computer Science 2024-08-15 Mark Benjamin , Gianluca Iaccarino

Due to their data-driven nature, Machine Learning (ML) models are susceptible to bias inherited from data, especially in classification problems where class and group imbalances are prevalent. Class imbalance (in the classification target)…

Machine Learning · Computer Science 2024-09-10 Emmanouil Panagiotou , Arjun Roy , Eirini Ntoutsi

The generation of data is a common approach to improve the performance of machine learning tasks, among which is the training of models for classification. In this paper, we present TAGAL, a collection of methods able to generate synthetic…

Machine Learning · Computer Science 2025-09-05 Benoît Ronval , Pierre Dupont , Siegfried Nijssen

Using a deep generative machine learning approach, we synthesise human activity participations and scheduling; i.e. the choices of what activities to participate in and when. Activity schedules are a core component of many applied…

Machine Learning · Computer Science 2025-10-03 Fred Shone , Tim Hillel

One of the limiting factors in training data-driven, rare-event prediction algorithms is the scarcity of the events of interest resulting in an extreme imbalance in the data. There have been many methods introduced in the literature for…

Machine Learning · Computer Science 2021-05-18 Yang Chen , Dustin J. Kempton , Azim Ahmadzadeh , Rafal A. Angryk

Generative Adversarial Networks (GANs) are gaining increasing attention as a means for synthesising data. So far much of this work has been applied to use cases outside of the data confidentiality domain with a common application being the…

Machine Learning · Computer Science 2021-12-06 Claire Little , Mark Elliot , Richard Allmendinger , Sahel Shariati Samani

This paper presents a comprehensive systematic review of generative models (GANs, VAEs, DMs, and LLMs) used to synthesize various medical data types, including imaging (dermoscopic, mammographic, ultrasound, CT, MRI, and X-ray), text,…

This research delves into the construction and utilization of synthetic datasets, specifically within the telematics sphere, leveraging OpenAI's powerful language model, ChatGPT. Synthetic datasets present an effective solution to…

Computers and Society · Computer Science 2023-06-27 Ryan Lingo

Synthetic data generation has been widely adopted in software testing, data privacy, imbalanced learning, and artificial intelligence explanation. In all such contexts, it is crucial to generate plausible data samples. A common assumption…

Artificial Intelligence · Computer Science 2024-10-16 Martina Cinquini , Fosca Giannotti , Riccardo Guidotti

Synthetic data generation, leveraging generative machine learning techniques, offers a promising approach to mitigating privacy concerns associated with real-world data usage. Synthetic data closely resembles real-world data while…

Machine Learning · Computer Science 2025-08-25 Weijie Niu , Alberto Huertas Celdran , Karoline Siarsky , Burkhard Stiller

We propose a unified framework for not only attributing synthetic speech to its source but also for detecting speech generated by synthesizers that were not encountered during training. This requires methods that move beyond simple…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Mohd Mujtaba Akhtar , Girish , Farhan Sheth , Muskaan Singh

Synthetic tabular data enables sharing and analysis of sensitive records, but its practical deployment requires balancing distributional fidelity, downstream utility, and privacy protection. We study a simple, model agnostic post processing…

Machine Learning · Computer Science 2026-02-09 David Yavo , Richard Khoury , Christophe Pere , Sadoune Ait Kaci Azzou

Strategies that include the generation of synthetic data are beginning to be viable as obtaining real data can be logistically complicated, very expensive or slow. Not only the capture of the data can lead to complications, but also its…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Paola Natalia Canas , Juan Diego Ortega , Marcos Nieto , Oihana Otaegui

Temporally indexed data are essential in a wide range of fields and of interest to machine learning researchers. Time series data, however, are often scarce or highly sensitive, which precludes the sharing of data between researchers and…

Machine Learning · Computer Science 2024-07-10 Alexander Nikitin , Letizia Iannucci , Samuel Kaski

The development of robust Document AI models has been constrained by limited access to high-quality, labeled datasets, primarily due to data privacy concerns, scarcity, and the high cost of manual annotation. Traditional methods of…

Computation and Language · Computer Science 2024-12-06 Amit Agarwal , Hitesh Patel , Priyaranjan Pattnayak , Srikant Panda , Bhargava Kumar , Tejaswini Kumar
‹ Prev 1 4 5 6 7 8 10 Next ›