English
Related papers

Related papers: Improving Equity in Health Modeling with GPT4-Turb…

200 papers

This paper explores the strategic use of modern synthetic data generation and advanced data perturbation techniques to enhance security, maintain analytical utility, and improve operational efficiency when managing large datasets, with a…

Cryptography and Security · Computer Science 2025-04-29 Anantha Sharma , Swetha Devabhaktuni , Eklove Mohan

Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify…

Machine Learning · Computer Science 2025-02-11 Tommaso Ferracci , Leonie Tabea Goldmann , Anton Hinel , Francesco Sanna Passino

The rapid advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have shown great potential in medical diagnostics, particularly in radiology, where datasets such as X-rays are paired with human-generated diagnostic…

Artificial Intelligence · Computer Science 2024-11-13 Francesco Chiumento , Mingming Liu

This paper examines biases in large language models (LLMs) when generating synthetic populations from responses to personality questionnaires. Using five LLMs, we first assess the representativeness and potential biases in the…

Computers and Society · Computer Science 2026-02-04 Jacopo Amidei , Gregorio Ferreira , Mario Muñoz Serrano , Rubén Nieto , Andreas Kaltenbrunner

Suicidal ideation detection is a vital research area that holds great potential for improving mental health support systems. However, the sensitivity surrounding suicide-related data poses challenges in accessing large-scale, annotated…

Computation and Language · Computer Science 2024-02-06 Hamideh Ghanadian , Isar Nejadgholi , Hussein Al Osman

In this work, we propose data augmentation via pairwise mixup across subgroups to improve group fairness. Many real-world applications of machine learning systems exhibit biases across certain groups due to under-representation or training…

Machine Learning · Statistics 2023-09-14 Madeline Navarro , Camille Little , Genevera I. Allen , Santiago Segarra

Data augmentations are useful in closing the sim-to-real domain gap when training on synthetic data. This is because they widen the training data distribution, thus encouraging the model to generalize better to other domains. Many image…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Bram Vanherle , Nick Michiels , Frank Van Reeth

This paper presents a comprehensive systematic review of generative models (GANs, VAEs, DMs, and LLMs) used to synthesize various medical data types, including imaging (dermoscopic, mammographic, ultrasound, CT, MRI, and X-ray), text,…

Developing and validating psychometric scales requires large samples, multiple testing phases, and substantial resources. Recent advances in Large Language Models (LLMs) enable the generation of synthetic participant data by prompting…

Human-Computer Interaction · Computer Science 2025-12-30 Enrico Cipriani , Pavel Okopnyi , Danilo Menicucci , Simone Grassini

Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We show that fully…

Methodology · Statistics 2026-05-12 Yichen Xu

Large language models (LLMs) have demonstrated remarkable in-context learning capabilities across diverse applications. In this work, we explore the effectiveness of LLMs for generating realistic synthetic tabular data, identifying key…

Machine Learning · Computer Science 2025-01-15 Jinhee Kim , Taesung Kim , Jaegul Choo

Large language models (LLMs), such as ChatGPT and Claude, have emerged as essential tools for information retrieval, often serving as alternatives to traditional search engines. However, ensuring that these models provide accurate and…

Computers and Society · Computer Science 2025-08-05 Uvini Balasuriya Mudiyanselage , Bharat Jayprakash , Kookjin Lee , K. Hazel Kwon

Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a…

Computation and Language · Computer Science 2024-10-10 Shiho Matta , Yin Jou Huang , Fei Cheng , Hirokazu Kiyomaru , Yugo Murawaki

Generative models have been showing potential for producing data in mass. This study explores the enhancement of clinical natural language processing performance by utilizing synthetic data generated from advanced language models. Promising…

Computation and Language · Computer Science 2024-03-29 Shan Chen , Jack Gallifant , Marco Guevara , Yanjun Gao , Majid Afshar , Timothy Miller , Dmitriy Dligach , Danielle S. Bitterman

Automatic hate speech detection using deep neural models is hampered by the scarcity of labeled datasets, leading to poor generalization. To mitigate this problem, generative AI has been utilized to generate large amounts of synthetic hate…

Computation and Language · Computer Science 2023-11-17 Sagi Pendzel , Tomer Wullach , Amir Adler , Einat Minkov

Despite the great advance of Multimodal Large Language Models (MLLMs) in both instruction dataset building and benchmarking, the independence of training and evaluation makes current MLLMs hard to further improve their capability under the…

Machine Learning · Computer Science 2023-09-12 Zhiyuan Zhao , Linke Ouyang , Bin Wang , Siyuan Huang , Pan Zhang , Xiaoyi Dong , Jiaqi Wang , Conghui He

Many machine learning models perform well when making predictions within the training data range, but often struggle when required to extrapolate beyond it. Symbolic regression (SR) using genetic programming (GP) can generate flexible…

Machine Learning · Computer Science 2025-12-01 Fitria Wulandari Ramlan , Colm O'Riordan , Gabriel Kronberger , James McDermott

Representational bias in large language models (LLMs) has predominantly been measured through single-response interactions and has focused on Global North-centric identities like race and gender. We expand on that research by conducting a…

Computation and Language · Computer Science 2025-08-07 Agrima Seth , Monojit Choudhary , Sunayana Sitaram , Kentaro Toyama , Aditya Vashistha , Kalika Bali

We investigate the usefulness of generative Large Language Models (LLMs) in generating training data for cross-encoder re-rankers in a novel direction: generating synthetic documents instead of synthetic queries. We introduce a new dataset,…

Information Retrieval · Computer Science 2023-05-04 Arian Askari , Mohammad Aliannejadi , Evangelos Kanoulas , Suzan Verberne

Personal thermal comfort models aim to predict an individual's thermal comfort response, instead of the average response of a large group. Recently, machine learning algorithms have proven to be having enormous potential as a candidate for…

Machine Learning · Computer Science 2022-11-22 Hari Prasanna Das , Costas J. Spanos