English

Identification Risks Evaluation of Partially Synthetic Data with the $\texttt{IdentificationRiskCalculation}$ R Package

Methodology 2021-04-07 v4 Applications

Abstract

We extend a general approach to evaluating identification risk of synthesized variables in partially synthetic data. For multiple continuous synthesized variables, we introduce the use of a radius rr in the construction of identification risk probability of each target record, and illustrate with working examples. We create the IdentificationRiskCalculation\texttt{IdentificationRiskCalculation} R package to aid researchers and data disseminators in performing these identification risks evaluation calculations. We demonstrate our methods through the R package with applications to a data sample from the Consumer Expenditure Surveys, and discuss the impacts on risk and data utility of 1) the choice of radius rr, 2) the choice of synthesized variables, and 3) the choice of number of synthetic datasets. We give recommendations for statistical agencies for synthesizing and evaluating identification risk of continuous variables.

Keywords

Cite

@article{arxiv.2006.01298,
  title  = {Identification Risks Evaluation of Partially Synthetic Data with the $\texttt{IdentificationRiskCalculation}$ R Package},
  author = {Ryan Hornby and Jingchen Hu},
  journal= {arXiv preprint arXiv:2006.01298},
  year   = {2021}
}

Comments

16 pages with 10 figures

R2 v1 2026-06-23T15:58:42.140Z