English
Related papers

Related papers: Best Practices and Lessons Learned on Synthetic Da…

200 papers

Ensuring safe adoption of AI tools in healthcare hinges on access to sufficient data for training, testing and validation. In response to privacy concerns and regulatory requirements, using synthetic data has been suggested. Synthetic data…

Sustainable AI is a subfield of AI for concerning developing and using AI systems in ways of aiming to reduce environmental impact and achieve sustainability. Sustainable AI is increasingly important given that training of and inference…

Artificial Intelligence · Computer Science 2025-01-14 Tao Xie , David Harel , Dezhi Ran , Zhenwen Li , Maoliang Li , Zhi Yang , Leye Wang , Xiang Chen , Ying Zhang , Wentao Zhang , Meng Li , Chen Zhang , Linyi Li , Assaf Marron

The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this study, we systematically investigate the feasibility of…

Sound · Computer Science 2025-12-22 Tingxiao Zhou , Leying Zhang , Zhengyang Chen , Yanmin Qian

Synthetic data has been advertised as a silver-bullet solution to privacy-preserving data publishing that addresses the shortcomings of traditional anonymisation techniques. The promise is that synthetic data drawn from generative models…

Machine Learning · Computer Science 2022-01-25 Theresa Stadler , Bristena Oprisanu , Carmela Troncoso

This paper demonstrates the potential of statistical disclosure control for protecting the data used to train recommender systems. Specifically, we use a synthetic data generation approach to hide specific information in the user-item…

Information Retrieval · Computer Science 2020-08-11 Manel Slokom , Martha Larson , Alan Hanjalic

With the proliferation of increasingly complicated Deep Learning architectures, data synthesis is a highly promising technique to address the demand of data-hungry models. However, reliably assessing the quality of a 'synthesiser' model's…

Machine Learning · Computer Science 2025-05-05 Julia A. Meister , Khuong An Nguyen

As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates the impact of the diversity of sources of synthetic data on fine-tuned large language…

Computation and Language · Computer Science 2026-04-29 Max Schaffelder , Albert Gatt

Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional…

Sound · Computer Science 2024-08-30 Tiantian Feng , Dimitrios Dimitriadis , Shrikanth Narayanan

Modern language model-based AI systems are remarkably powerful, yet their capabilities remain fundamentally capped by their human creators in three key ways. First, although a model's weights can be updated via fine-tuning, acquiring new…

Artificial Intelligence · Computer Science 2026-03-20 Zitong Yang

2023 was the year the world woke up to generative AI, and 2024 is the year policymakers are responding more firmly. Importantly, this policy momentum is taking place alongside real world creation and distribution of synthetic media. Social…

Computers and Society · Computer Science 2024-07-22 Claire R. Leibowicz , Christian H. Cardona

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Yuanchen Fei , Yude Zou , Zejian Kang , Ming Li , Jiaying Zhou , Xiangru Huang

The use of synthetic data has emerged as an essential tool in Magnetic Resonance Spectroscopy (MRS) research and applications, providing advantages for optimization of acquisition, software validation, deep learning applications, and…

The unparalleled success of artificial intelligence (AI) in the technology sector has catalyzed an enormous amount of research in the scientific community. It has proven to be a powerful tool, but as with any rapidly developing field, the…

Machine Learning · Computer Science 2022-10-07 M. R. Carbone

This study leverages synthetic data as a validation set to reduce overfitting and ease the selection of the best model in AI development. While synthetic data have been used for augmenting the training set, we find that synthetic data can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Qixin Hu , Alan Yuille , Zongwei Zhou

Synthetic data is a useful resource for algorithmic research. It allows for the evaluation of systems under a range of conditions that might be difficult to achieve in real world settings. In recommender systems, the use of synthetic data…

Information Retrieval · Computer Science 2024-09-24 Elena Stefancova , Cassidy All , Joshua Paup , Martin Homola , Nicholas Mattei , Robin Burke

The development of robust AI models relies heavily on the quality and variety of training data available. In fields where data scarcity is prevalent, synthetic data generation offers a vital solution. In this paper, we introduce a novel…

Computation and Language · Computer Science 2024-06-19 Elin Törnquist , Robert Alexander Caulk

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Vahid Noroozi , Zhehuai Chen , Somshubra Majumdar , Steve Huang , Jagadeesh Balam , Boris Ginsburg

Individual-level data (microdata) that characterizes a population, is essential for studying many real-world problems. However, acquiring such data is not straightforward due to cost and privacy constraints, and access is often limited to…

Machine Learning · Computer Science 2022-12-13 Angeela Acharya , Siddhartha Sikdar , Sanmay Das , Huzefa Rangwala

Collecting high quality conversational data can be very expensive for most applications and infeasible for others due to privacy, ethical, or similar concerns. A promising direction to tackle this problem is to generate synthetic dialogues…

Computation and Language · Computer Science 2023-02-20 Maximillian Chen , Alexandros Papangelis , Chenyang Tao , Seokhwan Kim , Andy Rosenbaum , Yang Liu , Zhou Yu , Dilek Hakkani-Tur

Data-centric AI is at the center of a fundamental shift in software engineering where machine learning becomes the new software, powered by big data and computing infrastructure. Here software engineering needs to be re-thought where data…

Machine Learning · Computer Science 2022-12-27 Steven Euijong Whang , Yuji Roh , Hwanjun Song , Jae-Gil Lee