English
Related papers

Related papers: BIGBOY1.2: Generating Realistic Synthetic Data for…

200 papers

Data imbalance in training data often leads to biased predictions from trained models, which in turn causes ethical and social issues. A straightforward solution is to carefully curate training data, but given the enormous scale of modern…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Moon Ye-Bin , Nam Hyeon-Woo , Wonseok Choi , Nayeong Kim , Suha Kwak , Tae-Hyun Oh

NLP researchers need more, higher-quality text datasets. Human-labeled datasets are expensive to collect, while datasets collected via automatic retrieval from the web such as WikiBio are noisy and can include undesired biases. Moreover,…

Computation and Language · Computer Science 2022-01-14 Ann Yuan , Daphne Ippolito , Vitaly Nikolaev , Chris Callison-Burch , Andy Coenen , Sebastian Gehrmann

Currently, there are many difficulties regarding the interoperability of medical data and related population data sources. These complications get in the way of the generation of high-quality data sets at city, region and national levels.…

Computers and Society · Computer Science 2023-10-13 Anna Andreychenko , Viktoriia Korzhuk , Stanislav Kondratenko , Polina Cheraneva

For many infectious disease outbreaks, the at-risk population changes their behavior in response to the outbreak severity, causing the transmission dynamics to change in real-time. Behavioral change is often ignored in epidemic modeling…

Methodology · Statistics 2023-10-25 Caitlin Ward , Rob Deardon , Alexandra M. Schmidt

Agent-based models (ABMs) simulate interactions between autonomous agents in constrained environments over time. ABMs are often used for modeling the spread of infectious diseases. In order to simulate disease outbreaks or other phenomena,…

Other Statistics · Statistics 2017-01-11 Shannon Gallagher , Lee Richardson , Samuel L. Ventura , William F. Eddy

Effective utilization of time series data is often constrained by the scarcity of data quantity that reflects complex dynamics, especially under the condition of distributional shifts. Existing datasets may not encompass the full range of…

Computational Engineering, Finance, and Science · Computer Science 2024-06-11 Haibei Zhu , Yousef El-Laham , Elizabeth Fons , Svitlana Vyetrenko

We consider a stochastic Susceptible-Exposed-Infected-Recovered (SEIR) epidemiological model with a contact rate that fluctuates seasonally. Through the use of a nonlinear, stochastic projection, we are able to analytically determine the…

Populations and Evolution · Quantitative Biology 2013-09-11 Eric Forgoston , Ira B. Schwartz

Evaluating time series attribution methods is difficult because real-world datasets rarely provide ground truth for which time points drive a prediction. A common workaround is to generate synthetic data where class-discriminating features…

Machine Learning · Computer Science 2026-03-10 Gregor Baer

This paper addresses statistical modelling and forecasting of key indicators describing the severity of a developing pandemic, using routinely reported daily counts of infections, hospitalizations, deaths (both in and out of hospital), and…

Despite significant recent progress in the area of Brain-Computer Interface (BCI), there are numerous shortcomings associated with collecting Electroencephalography (EEG) signals in real-world environments. These include, but are not…

Quantitative Methods · Quantitative Biology 2019-10-14 Nik Khadijah Nik Aznan , Amir Atapour-Abarghouei , Stephen Bonner , Jason Connolly , Noura Al Moubayed , Toby Breckon

Many ground-breaking advancements in machine learning can be attributed to the availability of a large volume of rich data. Unfortunately, many large-scale datasets are highly sensitive, such as healthcare data, and are not widely available…

Machine Learning · Computer Science 2020-12-09 James Jordon , Alan Wilson , Mihaela van der Schaar

Generative models have been found effective for data synthesis due to their ability to capture complex underlying data distributions. The quality of generated data from these models is commonly evaluated by visual inspection for image…

Machine Learning · Computer Science 2022-10-18 Emily Muller , Xu Zheng , Jer Hayes

Synthetic Data Generation (SDG), leveraging Large Language Models (LLMs), has recently been recognized and broadly adopted as an effective approach to improve the performance of smaller but more resource and compute efficient LLMs through…

Machine Learning · Computer Science 2026-03-25 Srideepika Jayaraman , Achille Fokoue , Dhaval Patel , Jayant Kalagnanam

Data synthesis for training large reasoning models offers a scalable alternative to limited, human-curated datasets, enabling the creation of high-quality data. However, existing approaches face several challenges: (i) indiscriminate…

Artificial Intelligence · Computer Science 2026-05-11 Yongxian Wei , Yilin Zhao , Zixuan Hu , Li Shen , Xinrui Chen , Runxi Cheng , Sinan Du , Hao Yu , Chun Yuan , Dian Li

Access to longitudinal, individual-level data on work-life balance and wellbeing is limited by privacy, ethical, and logistical constraints. This poses challenges for reproducible research, methodological benchmarking, and education in…

Machine Learning · Computer Science 2025-12-30 Wafaa El Husseini

Evaluating the performance of machine learning models on diverse and underrepresented subgroups is essential for ensuring fairness and reliability in real-world applications. However, accurately assessing model performance becomes…

Machine Learning · Computer Science 2023-10-26 Boris van Breugel , Nabeel Seedat , Fergus Imrie , Mihaela van der Schaar

*Data Synthesis* is a promising way to train a small model with very little labeled data. One approach for data synthesis is to leverage the rich knowledge from large language models to synthesize pseudo training examples for small models,…

Computation and Language · Computer Science 2023-10-23 Ruida Wang , Wangchunshu Zhou , Mrinmaya Sachan

We release SVIRO, a synthetic dataset for sceneries in the passenger compartment of ten different vehicles, in order to analyze machine learning-based approaches for their generalization capacities and reliability when trained on a limited…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Steve Dias Da Cruz , Oliver Wasenmüller , Hans-Peter Beise , Thomas Stifter , Didier Stricker

The lack of freely available standardized datasets represents an aggravating factor during the development and testing the performance of novel computational techniques in exposure assessment and dosimetry research. This hinders progress as…

Medical Physics · Physics 2023-05-04 Ante Kapetanovic , Dragan Poljak , Kun Li

$\textbf{Background:}$ At the onset of a pandemic, such as COVID-19, data with proper labeling/attributes corresponding to the new disease might be unavailable or sparse. Machine Learning (ML) models trained with the available data, which…

‹ Prev 1 8 9 10 Next ›