English

Towards Inference-Aware Privacy Guidance for Data Preparation

Databases 2026-07-18 v1

Abstract

Data preparation often begins with sensitive data and produces a releasable artifact for analysis, sharing, or model training. Existing workflows are primarily guided by utility: a curator drops attributes, coarsens values, filters populations, and suppresses tuples until the resulting dataset appears useful and safe. Privacy, when considered, is usually evaluated only on the final release. We propose privacy-aware data preparation as an interactive guidance problem. We model a preparation plan as a sequence of deterministic curation operators and ask how each step changes the evidence available to an observer with prior knowledge about a target. Our semantics is based on compatibility sets, which capture the source tuples still plausible for the target after a released representation is observed. This view separates operators that remove evidence from those that remove ambiguity, explains why privacy effects can be non-monotone, and supports prefix-level feedback under a disclosure budget. The result is an inference-aware foundation for guiding curators throughout data preparation, rather than judging privacy only after the final artifact is produced. We conclude by identifying the key challenges in building interactive, inference-aware data preparation systems.

Cite

@article{arxiv.2607.16710,
  title  = {Towards Inference-Aware Privacy Guidance for Data Preparation},
  author = {Vishal Chakraborty and Felix Naumann},
  journal= {arXiv preprint arXiv:2607.16710},
  year   = {2026}
}

Comments

VLDB 2026 Workshop:15th International Workshop on Quality in Databases (QDB 2026)