中文
相关论文

相关论文: Risk-Efficient Bayesian Data Synthesis for Privacy…

200 篇论文

Differentially private (DP) synthetic data generation is a practical method for improving access to data as a means to encourage productive partnerships. One issue inherent to DP is that the "privacy budget" is generally "spent" evenly…

机器学习 · 计算机科学 2022-08-11 Lucas Rosenblatt , Joshua Allen , Julia Stoyanovich

Differential privacy is a mathematical concept that provides an information-theoretic security guarantee. While differential privacy has emerged as a de facto standard for guaranteeing privacy in data sharing, the known mechanisms to…

密码学与安全 · 计算机科学 2024-03-26 March Boedihardjo , Thomas Strohmer , Roman Vershynin

Smartwatch health sensor data are increasingly utilized in smart health applications and patient monitoring, including stress detection. However, such medical data often comprise sensitive personal information and are resource-intensive to…

机器学习 · 计算机科学 2026-02-03 Lucas Lange , Nils Wenzlitschke , Erhard Rahm

Ratio statistics--such as relative risk and odds ratios--play a central role in hypothesis testing, model evaluation, and decision-making across many areas of machine learning, including causal inference and fairness analysis. However,…

机器学习 · 统计学 2025-05-28 Tomer Shoham , Katrina Ligettt

Privacy Preserving Synthetic Data Generation (PP-SDG) has emerged to produce synthetic datasets from personal data while maintaining privacy and utility. Differential privacy (DP) is the property of a PP-SDG mechanism that establishes how…

Synthetic data generation (SDG) has become increasingly popular as a privacy-enhancing technology. It aims to maintain important statistical properties of its underlying training data, while excluding any personally identifiable…

密码学与安全 · 计算机科学 2024-02-13 Steven Golob , Sikha Pentyala , Anuar Maratkhan , Martine De Cock

The digitization of medical records ushered in a new era of big data to clinical science, and with it the possibility that data could be shared, to multiply insights beyond what investigators could abstract from paper records. The need to…

机器学习 · 计算机科学 2021-06-03 Ofer Mendelevitch , Michael D. Lesh

Synthetic data generation is an appealing tool for augmenting and enriching datasets, playing a crucial role in advancing artificial intelligence (AI) and machine learning (ML). Not only does synthetic data help build robust AI/ML datasets…

系统与控制 · 电气工程与系统科学 2026-03-20 José Pulido , Francesc Wilhelmi , Sergio Fortes , Alfonso Fernández-Durán , Lorenzo Galati Giordano , Raquel Barco

Synthetic tabular data is essential for machine learning workflows, especially for expanding small or imbalanced datasets and enabling privacy-preserving data sharing. However, state-of-the-art generative models (GANs, VAEs, diffusion…

机器学习 · 计算机科学 2025-07-24 Jessup Byun , Xiaofeng Lin , Joshua Ward , Guang Cheng

Benchmarking is crucial for evaluating a DBMS, yet existing benchmarks often fail to reflect the varied nature of user workloads. As a result, there is increasing momentum toward creating databases that incorporate real-world user data to…

数据库 · 计算机科学 2025-04-11 Yunqing Ge , Jianbin Qin , Shuyuan Zheng , Yongrui Zhong , Bo Tang , Yu-Xuan Qiu , Rui Mao , Ye Yuan , Makoto Onizuka , Chuan Xiao

Differential privacy (DP) provides a mathematical guarantee limiting what an adversary can learn about any individual from released data. However, achieving this protection typically requires adding noise, and noise can accumulate when many…

机器学习 · 计算机科学 2026-02-12 Amir Asiaee , Chao Yan , Zachary B. Abrams , Bradley A. Malin

Privacy-preserving synthetic data offers a promising solution to harness segregated data in high-stakes domains where information is compartmentalized for regulatory, privacy, or institutional reasons. This survey provides a comprehensive…

密码学与安全 · 计算机科学 2025-03-28 Viktor Schlegel , Anil A Bharath , Zilong Zhao , Kevin Yee

Bayesian networks (BN) are probabilistic graphical models that enable efficient knowledge representation and inference. These have proven effective across diverse domains, including healthcare, bioinformatics and economics. The structure…

机器学习 · 计算机科学 2026-02-24 Niccolò Rocchi , Fabio Stella , Cassio de Campos

As data-driven and AI-based decision making gains widespread adoption across disciplines, it is crucial that both data privacy and decision fairness are appropriately addressed. Although differential privacy (DP) provides a robust framework…

机器学习 · 计算机科学 2025-10-21 Spencer Giddens , Xiaon Lang , Fang Liu

Synthetic control is a causal inference tool used to estimate the treatment effects of an intervention by creating synthetic counterfactual data. This approach combines measurements from other similar observations (i.e., donor pool ) to…

机器学习 · 计算机科学 2023-03-27 Saeyoung Rho , Rachel Cummings , Vishal Misra

Micro and survey datasets often contain private information about individuals, like their health status, income or political preferences. Previous studies have shown that, even after data anonymization, a malicious intruder could still be…

应用统计 · 统计学 2024-08-26 Marco Battiston , Lorenzo Rimella

Machine learning systems require representations of the real world for training and testing - they require data, and lots of it. Collecting data at scale has logistical and ethical challenges, and synthetic data promises a solution to these…

计算机与社会 · 计算机科学 2024-05-06 Cedric Deslandes Whitney , Justin Norman

The Synthetic Minority Over-sampling Technique (SMOTE) is one of the most widely used methods for addressing class imbalance and generating synthetic data. Despite its popularity, little attention has been paid to its privacy implications;…

密码学与安全 · 计算机科学 2026-03-03 Georgi Ganev , Reza Nazari , Rees Davison , Amir Dizche , Xinmin Wu , Ralph Abbey , Jorge Silva , Emiliano De Cristofaro

This paper presents a privacy-preserving event detection scheme based on measurements made by a network of sensors. A diameter-like decision statistic made up of the marginal types of the measurements observed by the sensors is employed.…

信息论 · 计算机科学 2025-05-06 Xiaoshan Wang , Tan F. Wong

Deep generative models are often trained on sensitive data, such as genetic sequences, health data, or more broadly, any copyrighted, licensed or protected content. This raises critical concerns around privacy-preserving synthetic data, and…