English
Related papers

Related papers: Building a large synthetic population from Austral…

200 papers

High-quality human mobility data is crucial for applications such as urban planning, transportation management, and public health, yet its collection is often hindered by privacy concerns and data scarcity-particularly in less-developed…

Social and Information Networks · Computer Science 2025-12-19 Yuan Yuan , Yuheng Zhang , Jingtao Ding , Yong Li

In recent years, as smart home systems have become more widespread, security concerns within these environments have become a growing threat. Currently, most smart home security solutions, such as anomaly detection and behavior prediction…

Artificial Intelligence · Computer Science 2025-02-03 Zhiyao Xu , Dan Zhao , Qingsong Zou , Jingyu Xiao , Yong Jiang , Zhenhui Yuan , Qing Li

Public-use microdata samples (PUMS) from the United States (US) Census Bureau on individuals have been available for decades. However, large increases in computing power and the greater availability of Big Data have dramatically increased…

In many simulation studies involving networks there is the need to rely on a sample network to perform the simulation experiments. In many cases, real network data is not available due to privacy concerns. In that case we can recourse to…

Social and Information Networks · Computer Science 2014-11-25 Hebert Pérez-Rosés , Francesc Sebé

Recently, powerful Large Language Models (LLMs) have become easily accessible to hundreds of millions of users world-wide. However, their strong capabilities and vast world knowledge do not come without associated privacy risks. In this…

Machine Learning · Computer Science 2024-11-05 Hanna Yukhymenko , Robin Staab , Mark Vero , Martin Vechev

Generating synthetic data through generative models is gaining interest in the ML community and beyond. In the past, synthetic data was often regarded as a means to private data release, but a surge of recent papers explore how its…

Machine Learning · Computer Science 2023-04-10 Boris van Breugel , Mihaela van der Schaar

Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and…

Artificial Intelligence · Computer Science 2026-04-01 Tim R. Davidson , Benoit Seguin , Enrico Bacis , Cesar Ilharco , Hamza Harkous

State-space models are commonly used to describe different forms of ecological data. We consider the case of count data with observation errors. For such data the system process is typically multi-dimensional consisting of coupled Markov…

Methodology · Statistics 2017-08-15 Axel Finke , Ruth King , Alexandros Beskos , Petros Dellaportas

The way LLM-based entities conceive of the relationship between AI and humans is an important topic for both cultural and safety reasons. When we examine this topic, what matters is not only the model itself but also the personas we…

Computation and Language · Computer Science 2026-02-27 Jiří Milička , Hana Bednářová

The U.S. Census Bureau provides an estimate of the true population as a supplement to the basic census numbers. This estimate is constructed from data in a post-censal survey. The overall procedure is referred to as dual system estimation.…

Applications · Statistics 2008-12-18 Lawrence Brown , Zhanyun Zhao

We present population synthesis modeling of the X-ray background with genetic algorithm - based optimization method. In our models the best fit could be achieved for lower values of high-energy exponential cut-off (~ 170 keV) and larger…

Astrophysics · Physics 2007-05-23 Alexander V. Halevin

Synthetic data has been widely applied in the real world recently. One typical example is the creation of synthetic data for privacy concerned datasets. In this scenario, synthetic data substitute the real data which contains the privacy…

Software Engineering · Computer Science 2023-12-12 Xiao Ling , Tim Menzies , Christopher Hazard , Jack Shu , Jacob Beel

Large language models (LLMs) have great potential for synthetic data generation. This work shows that useful data can be synthetically generated even for tasks that cannot be solved directly by LLMs: for problems with structured outputs, it…

Computation and Language · Computer Science 2023-10-31 Martin Josifoski , Marija Sakota , Maxime Peyrard , Robert West

Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset…

Machine Learning · Computer Science 2022-11-08 Gift Khangamwa , Terence L. van Zyl , Clint J. van Alten

Many aspects of the evolution of stars, and in particular the evolution of binary stars, remain beyond our ability to model them in detail. Instead, we rely on observations to guide our often phenomenological models and pin down uncertain…

Solar and Stellar Astrophysics · Physics 2018-08-22 Robert G. Izzard , Ghina M. Halabi

Dense crowd counting aims to predict thousands of human instances from an image, by calculating integrals of a density map over image pixels. Existing approaches mainly suffer from the extreme density variances. Such density pattern shift…

Computer Vision and Pattern Recognition · Computer Science 2019-08-09 Chenfeng Xu , Kai Qiu , Jianlong Fu , Song Bai , Yongchao Xu , Xiang Bai

With the dawn of the Big Data era, data sets are growing rapidly. Data is streaming from everywhere - from cameras, mobile phones, cars, and other electronic devices. Clustering streaming data is a very challenging problem. Unlike the…

Machine Learning · Computer Science 2019-02-08 Shlomo Bugdary , Shay Maymon

We introduce a modified spatial $\Lambda$-Fleming-Viot process to model the ancestry of individuals in a population occupying a continuous spatial habitat divided into two areas by a sharp discontinuity of the dispersal rate and effective…

Probability · Mathematics 2023-06-14 Raphael Forien , Harald Ringbauer , Graham Coop

Collecting high quality conversational data can be very expensive for most applications and infeasible for others due to privacy, ethical, or similar concerns. A promising direction to tackle this problem is to generate synthetic dialogues…

Computation and Language · Computer Science 2023-02-20 Maximillian Chen , Alexandros Papangelis , Chenyang Tao , Seokhwan Kim , Andy Rosenbaum , Yang Liu , Zhou Yu , Dilek Hakkani-Tur

Recent advances in deep learning methods have increased the performance of face detection and recognition systems. The accuracy of these models relies on the range of variation provided in the training data. Creating a dataset that…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Shubhajit Basak , Hossein Javidnia , Faisal Khan , Rachel McDonnell , Michael Schukat