English
Related papers

Related papers: Developing synthetic microdata through machine lea…

200 papers

The increased use of differential privacy (DP) has allowed the sharing of large amounts of data while reducing the risk of disclosure of sensitive information at the individual level. However, the noise introduced by DP methods makes…

Methodology · Statistics 2026-04-29 Jordan Awan , Xi Chen , Roberto Molinari

Privacy poses a significant obstacle to the progress of learning analytics (LA), presenting challenges like inadequate anonymization and data misuse that current solutions struggle to address. Synthetic data emerges as a potential remedy,…

Cryptography and Security · Computer Science 2024-01-17 Qinyi Liu , Mohammad Khalil , Ronas Shakya , Jelena Jovanovic

This paper introduces SynDiffix, a mechanism for generating statistically accurate, anonymous synthetic data for structured data. Recent open source and commercial systems use Generative Adversarial Networks or Transformed Auto Encoders to…

Cryptography and Security · Computer Science 2023-11-17 Paul Francis , Cristian Berneanu , Edon Gashi

Machine learning applications are becoming increasingly pervasive in our society. Since these decision-making systems rely on data-driven learning, risk is that they will systematically spread the bias embedded in data. In this paper, we…

Machine Learning · Statistics 2023-02-09 Alessandro Castelnovo , Riccardo Crupi , Nicole Inverardi , Daniele Regoli , Andrea Cosentini

In this paper, we suggest a systematic approach for developing socio-technical assessment for hiring ADS. We suggest using a matrix to expose underlying assumptions rooted in pseudoscientific essentialized understandings of human nature and…

Computers and Society · Computer Science 2022-05-13 Mona Sloane , Emanuel Moss , Rumman Chowdhury

Good training data is a prerequisite to develop useful ML applications. However, in many domains existing data sets cannot be shared due to privacy regulations (e.g., from medical studies). This work investigates a simple yet unconventional…

Synthetic data generation has been a growing area of research in recent years. However, its potential applications in serious games have not been thoroughly explored. Advances in this field could anticipate data modelling and analysis, as…

Computers and Society · Computer Science 2024-01-30 Jaime Pérez , Mario Castro , Edmond Awad , Gregorio López

Data augmentation is rapidly gaining attention in machine learning. Synthetic data can be generated by simple transformations or through the data distribution. In the latter case, the main challenge is to estimate the label associated to…

Machine Learning · Computer Science 2019-03-26 Maria Perez-Ortiz , Peter Tino , Rafal Mantiuk , Cesar Hervas-Martinez

Introduction: The amount of data generated by original research is growing exponentially. Publicly releasing them is recommended to comply with the Open Science principles. However, data collected from human participants cannot be released…

Machine Learning · Statistics 2023-10-11 Rémy Chapelle , Bruno Falissard

Generative models producing synthetic data are meant to provide a privacy-friendly approach to releasing data. However, their privacy guarantees are only considered robust when models satisfy Differential Privacy (DP). Alas, this is not a…

Cryptography and Security · Computer Science 2025-05-09 Georgi Ganev , Emiliano De Cristofaro

Digital footprints (records of individuals' interactions with digital systems) are essential for studying behavior, developing personalized applications, and training machine learning models. However, research in this area is often hindered…

Computation and Language · Computer Science 2026-03-13 Minjia Wang , Yunfeng Wang , Xiao Ma , Dexin Lv , Qifan Guo , Lynn Zheng , Benliang Wang , Lei Wang , Jiannan Li , Yongwei Xing , David Xu , Zheng Sun

Understanding urban mobility patterns and analyzing how people move around cities helps improve the overall quality of life and supports the development of more livable, efficient, and sustainable urban areas. A challenging aspect of this…

Computers and Society · Computer Science 2024-09-05 Prabin Bhandari , Antonios Anastasopoulos , Dieter Pfoser

Face recognition applications have grown in parallel with the size of datasets, complexity of deep learning models and computational power. However, while deep learning models evolve to become more capable and computational power keeps…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Pedro C. Neto , Rafael M. Mamede , Carolina Albuquerque , Tiago Gonçalves , Ana F. Sequeira

This paper studies the feasibility of synthetic data generation for mission-critical applications. The emphasis is on synthetic data generation for anomalous detection in complex social networks. In particular, the development of a…

Social and Information Networks · Computer Science 2020-10-27 Andreea Sistrunk , Vanessa Cedeno , Subhodip Biswas

Social science often relies on surveys of households and individuals. Dozens of such surveys are regularly administered by the U.S. government. However, they field independent, unconnected samples with specialized questions, limiting…

The public availability of collections containing user preferences is of vital importance for performing offline evaluations in the field of recommender systems. However, the number of rating datasets is limited because of the costs…

Information Retrieval · Computer Science 2019-09-04 Diego Monti , Giuseppe Rizzo , Maurizio Morisio

There is a known tension between the need to analyze personal data to drive business and privacy concerns. Many data protection regulations, including the EU General Data Protection Regulation (GDPR) and the California Consumer Protection…

Cryptography and Security · Computer Science 2022-02-02 Abigail Goldsteen , Gilad Ezov , Ron Shmelkin , Micha Moffie , Ariel Farkash

Recent work has shown the benefits of synthetic data for use in computer vision, with applications ranging from autonomous driving to face landmark detection and reconstruction. There are a number of benefits of using synthetic data from…

Computer Vision and Pattern Recognition · Computer Science 2023-01-04 Charlie Hewitt , Tadas Baltrušaitis , Erroll Wood , Lohit Petikam , Louis Florentin , Hanz Cuevas Velasquez

We propose a new system identification method, called Sign-Perturbed Sums (SPS), for constructing non-asymptotic confidence regions under mild statistical assumptions. SPS is introduced for linear regression models, including but not…

Signal Processing · Electrical Eng. & Systems 2018-07-24 Balázs Cs. Csáji , Marco C. Campi , Erik Weyer

Personal data centralization among dominant platform providers including search engines, social networking services, and e-commerce has created siloed ecosystems that restrict user sovereignty, thereby impeding data use across services.…

Artificial Intelligence · Computer Science 2026-02-11 Akinori Maeda , Yuto Sekiya , Sota Sugimura , Tomoya Asai , Yu Tsuda , Kohei Ikeda , Hiroshi Fujii , Kohei Watanabe