English
Related papers

Related papers: Generating Synthetic Ground Truth Distributions fo…

200 papers

Photometric redshift estimation algorithms are often based on representative data from observational campaigns. Data-driven methods of this type are subject to a number of potential deficiencies, such as sample bias and incompleteness.…

Cosmology and Nongalactic Astrophysics · Physics 2022-07-06 Nesar Ramachandra , Jonás Chaves-Montero , Alex Alarcon , Arindam Fadikar , Salman Habib , Katrin Heitmann

Recent studies have highlighted the benefits of generating multiple synthetic datasets for supervised learning, from increased accuracy to more effective model selection and uncertainty estimation. These benefits have clear empirical…

Machine Learning · Computer Science 2025-04-28 Ossi Räisä , Antti Honkela

Generative Adversarial Networks (GANs) have been used to model the underlying probability distribution of sample based datasets. GANs are notoriuos for training difficulties and their dependence on arbitrary hyperparameters. One recent…

Machine Learning · Computer Science 2019-10-03 Thomas Pinetz , Daniel Soukup , Thomas Pock

The use of synthetic data in machine learning applications and research offers many benefits, including performance improvements through data augmentation, privacy preservation of original samples, and reliable method assessment with fully…

Machine Learning · Computer Science 2026-04-13 Joanna Komorniczak

Synthetic network traffic generation has emerged as a promising alternative for various data-driven applications in the networking domain. It enables the creation of synthetic data that preserves real-world characteristics while addressing…

Networking and Internet Architecture · Computer Science 2026-04-16 Nirhoshan Sivaroopan , Kaushitha Silva , Chamara Madarasingha , Thilini Dahanayaka , Guillaume Jourjon , Anura Jayasumana , Kanchana Thilakarathna

Networks are widely used in science and technology to represent relationships between entities, such as social or ecological links between organisms, enzymatic interactions in metabolic systems, or computer infrastructure. Statistical…

Discrete Mathematics · Computer Science 2012-07-19 Alexander Gutfraind , Lauren Ancel Meyers , Ilya Safro

This article describes techniques employed in the production of a synthetic dataset of driver telematics emulated from a similar real insurance dataset. The synthetic dataset generated has 100,000 policies that included observations about…

Machine Learning · Statistics 2021-02-02 Banghee So , Jean-Philippe Boucher , Emiliano A. Valdez

Reliable probability estimation is of crucial importance in many real-world applications where there is inherent (aleatoric) uncertainty. Probability-estimation models are trained on observed outcomes (e.g. whether it has rained or not, or…

Many data stewards collect confidential data that include fine geography. When sharing these data with others, data stewards strive to disseminate data that are informative for a wide range of spatial and non-spatial analyses while…

Methodology · Statistics 2016-02-16 Harrison Quick , Scott H. Holan , Christopher K. Wikle , Jerome P. Reiter

Trajectory data mining is crucial for smart city management. However, collecting large-scale trajectory datasets is challenging due to factors such as commercial conflicts and privacy regulations. Therefore, we urgently need trajectory…

Machine Learning · Computer Science 2025-02-04 Jingyuan Wang , Yujing Lin , Yudong Li

Trajectory prediction is a fundamental and challenging task for numerous applications, such as autonomous driving and intelligent robots. Currently, most of existing work treat the pedestrian trajectory as a series of fixed two-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2021-03-17 Pei Lv , Hui Wei , Tianxin Gu , Yuzhen Zhang , Xiaoheng Jiang , Bing Zhou , Mingliang Xu

We propose conditional flows of the maximum mean discrepancy (MMD) with the negative distance kernel for posterior sampling and conditional generative modeling. This MMD, which is also known as energy distance, has several advantageous…

Score-based generative models are shown to achieve remarkable empirical performances in various applications such as image generation and audio synthesis. However, a theoretical understanding of score-based diffusion models is still…

Machine Learning · Computer Science 2022-12-14 Dohyun Kwon , Ying Fan , Kangwook Lee

Existing lane-level simulation road network generation is labor-intensive, resource-demanding, and costly due to the need for large-scale data collection and manual post-editing. To overcome these limitations, we propose automatically…

Multimedia · Computer Science 2025-09-04 Liang Xie , Wenke Huang

User mobility modeling serves a crucial role in analysis and optimization of contemporary wireless networks. Typical stochastic mobility models, e.g., random waypoint model and Gauss Markov model, can hardly capture the distribution…

Artificial Intelligence · Computer Science 2024-07-30 Zhenyu Tao , Wei Xu , Xiaohu You

A novel variational inference based resampling framework is proposed to evaluate the robustness and generalization capability of deep learning models with respect to distribution shift. We use Auto Encoding Variational Bayes to find a…

Machine Learning · Computer Science 2019-10-29 Xudong Sun , Alexej Gossmann , Yu Wang , Bernd Bischl

This work presents several expected generalization error bounds based on the Wasserstein distance. More specifically, it introduces full-dataset, single-letter, and random-subset bounds, and their analogues in the randomized subsample…

Machine Learning · Statistics 2022-03-29 Borja Rodríguez-Gálvez , Germán Bassi , Ragnar Thobaben , Mikael Skoglund

When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our…

Methodology · Statistics 2024-04-30 Shirley Mathur , Yajuan Si , Jerome P. Reiter

Despite several works that succeed in generating synthetic data with differential privacy (DP) guarantees, they are inadequate for generating high-quality synthetic data when the input data has missing values. In this work, we formalize the…

Databases · Computer Science 2025-11-06 Shubhankar Mohapatra , Jianqiao Zong , Florian Kerschbaum , Xi He

Dataset condensation constructs compact synthetic datasets that retain the training utility of large real-world datasets, enabling efficient model development and potentially supporting downstream research in governed domains such as…

Machine Learning · Computer Science 2026-04-24 Pafue Christy Nganjimi , Andrew Soltan , Danielle Belgrave , Lei Clifton , David Clifton , Anshul Thakur