中文
相关论文

相关论文: Copula-based transferable models for synthetic pop…

200 篇论文

This paper proposes a new method to generate synthetic data sets based on copula models. Our goal is to produce surrogate data resembling real data in terms of marginal and joint distributions. We present a complete and reliable algorithm…

机器学习 · 计算机科学 2022-04-01 Regis Houssou , Mihai-Cezar Augustin , Efstratios Rappos , Vivien Bonvin , Stephan Robert-Nicoud

The ability to generate high-fidelity synthetic data is crucial when available (real) data is limited or where privacy and data protection standards allow only for limited use of the given data, e.g., in medical and financial data-sets.…

机器学习 · 统计学 2021-01-05 Sanket Kamthe , Samuel Assefa , Marc Deisenroth

Synthetic population generation is the process of combining multiple socioeconomic and demographic datasets from different sources and/or granularity levels, and downscaling them to an individual level. Although it is a fundamental step for…

机器学习 · 计算机科学 2019-11-12 Colin Wan , Zheng Li , Alicia Guo , Yue Zhao

A synthetic dataset is a data object that is generated programmatically, and it may be valuable to creating a single dataset from multiple sources when direct collection is difficult or costly. Although it is a fundamental step for many…

应用统计 · 统计学 2020-09-22 Zheng Li , Yue Zhao , Jialin Fu

We introduce a constraint-programming framework for generating synthetic populations that reproduce target statistics with high precision while enforcing full individual consistency. Unlike data-driven approaches that infer distributions…

机器学习 · 统计学 2025-12-09 Thierry Petit , Arnault Pachot

Individual-level data (microdata) that characterizes a population, is essential for studying many real-world problems. However, acquiring such data is not straightforward due to cost and privacy constraints, and access is often limited to…

机器学习 · 计算机科学 2022-12-13 Angeela Acharya , Siddhartha Sikdar , Sanmay Das , Huzefa Rangwala

Exploring the dependence between covariates across distributions is crucial for many applications. Copulas serve as a powerful tool for modeling joint variable dependencies and have been effectively applied in various practical contexts due…

机器学习 · 统计学 2026-04-09 Sumin Wang , Chenxian Huang , Yongdao Zhou , Min-Qian Liu

Copulas are a powerful tool for modeling multivariate distributions as they allow to separately estimate the univariate marginal distributions and the joint dependency structure. However, known parametric copulas offer limited flexibility…

机器学习 · 统计学 2021-11-11 Tim Janke , Mohamed Ghanmi , Florian Steinke

Synthetic population is an increasingly important material used in numerous areas such as urban and transportation analysis. Traditional methods such as iterative proportional fitting (IPF) is not capable of generating high-quality data…

计算机与社会 · 计算机科学 2025-08-14 Hai Yang , Hongying Wu , Linfei Yuan , Xiyuan Ren , Joseph Y. J. Chow , Jinqin Gao , Kaan Ozbay

Population synthesis is a critical task that involves generating synthetic yet realistic representations of populations. It is a fundamental problem in agent-based modeling (ABM), which has become the standard to analyze intelligent…

机器学习 · 计算机科学 2025-08-14 Min Tang , Peng Lu , Qing Feng

Reliable machine learning and statistical analysis rely on diverse, well-distributed training data. However, real-world datasets are often limited in size and exhibit underrepresentation across key subpopulations, leading to biased…

统计方法学 · 统计学 2025-07-15 Xinyu Tian , Xiaotong Shen

Can we improve machine-learning (ML) emulators with synthetic data? If data are scarce or expensive to source and a physical model is available, statistically generated data may be useful for augmenting training sets cheaply. Here we…

机器学习 · 计算机科学 2021-09-28 David Meyer , Thomas Nagler , Robin J. Hogan

Synthetic contact networks are useful for modeling epidemic spread and social transmission, but data to infer realistic contact patterns that take account of assortative connections at the geographic and economic levels is limited. We…

社会与信息网络 · 计算机科学 2024-06-24 Alexander Y. Tulchinsky , Fardad Haghpanah , Alisa Hamilton , Nodar Kipshidze , Eili Y. Klein

The Copula is widely used to describe the relationship between the marginal distribution and joint distribution of random variables. The estimation of high-dimensional Copula is difficult, and most existing solutions rely either on…

机器学习 · 计算机科学 2022-11-02 Zhi Zeng , Ting Wang

In this article, we develop fully Bayesian, copula-based, spatial-statistical models for large, noisy, incomplete, and non-Gaussian spatial data. Our approach includes novel constructions of copulas that accommodate a spatial-random-effects…

统计方法学 · 统计学 2025-11-05 Alan Pearse , David Gunawan , Noel Cressie

By sampling from the latent space of an autoencoder and decoding the latent space samples to the original data space, any autoencoder can simply be turned into a generative model. For this to work, it is necessary to model the autoencoder's…

机器学习 · 统计学 2023-09-19 Maximilian Coblenz , Oliver Grothe , Fabian Kächele

It is increasingly important to generate synthetic populations with explicit coordinates rather than coarse geographic areas, yet no established methods exist to achieve this. One reason is that latitude and longitude differ from other…

机器学习 · 计算机科学 2025-10-14 Jacopo Lenti , Lorenzo Costantini , Ariadna Fosch , Anna Monticelli , David Scala , Marco Pangallo

Key to effective generic, or "black-box", variational inference is the selection of an approximation to the target density that balances accuracy and speed. Copula models are promising options, but calibration of the approximation can be…

统计方法学 · 统计学 2022-07-01 Michael Stanley Smith , Rubén Loaiza-Maya

To advance Educational Data Mining (EDM) within strict privacy-protecting regulatory frameworks, researchers must develop methods that enable data-driven analysis while protecting sensitive student information. Synthetic data generation is…

机器学习 · 计算机科学 2026-04-07 Gabriel Diaz Ramos , Lorenzo Luzi , Debshila Basu Mallick , Richard Baraniuk

Quantifying uncertainty in future climate projections is hindered by the prohibitive computational cost of running physical climate models, which severely limits the availability of training data. We propose a data-efficient framework for…

统计方法学 · 统计学 2026-02-27 Johannes Brachem , Paul F. V. Wiemann , Matthias Katzfuss
‹ 上一页 1 2 3 10 下一页 ›