English

Bayesian Advantage of Re-Identification Attack in the Shuffle Model

Cryptography and Security 2025-11-06 v1

Abstract

The shuffle model, which anonymizes data by randomly permuting user messages, has been widely adopted in both cryptography and differential privacy. In this work, we present the first systematic study of the Bayesian advantage in re-identifying a user's message under the shuffle model. We begin with a basic setting: one sample is drawn from a distribution PP, and n1n - 1 samples are drawn from a distribution QQ, after which all nn samples are randomly shuffled. We define βn(P,Q)\beta_n(P, Q) as the success probability of a Bayes-optimal adversary in identifying the sample from PP, and define the additive and multiplicative Bayesian advantages as Advn+(P,Q)=βn(P,Q)1n\mathsf{Adv}_n^{+}(P, Q) = \beta_n(P,Q) - \frac{1}{n} and Advn×(P,Q)=nβn(P,Q)\mathsf{Adv}_n^{\times}(P, Q) = n \cdot \beta_n(P,Q), respectively. We derive exact analytical expressions and asymptotic characterizations of βn(P,Q)\beta_n(P, Q), along with evaluations in several representative scenarios. Furthermore, we establish (nearly) tight mutual bounds between the additive Bayesian advantage and the total variation distance. Finally, we extend our analysis beyond the basic setting and present, for the first time, an upper bound on the success probability of Bayesian attacks in shuffle differential privacy. Specifically, when the outputs of nn users -- each processed through an ε\varepsilon-differentially private local randomizer -- are shuffled, the probability that an attacker successfully re-identifies any target user's message is at most eε/ne^{\varepsilon}/n.

Keywords

Cite

@article{arxiv.2511.03213,
  title  = {Bayesian Advantage of Re-Identification Attack in the Shuffle Model},
  author = {Pengcheng Su and Haibo Cheng and Ping Wang},
  journal= {arXiv preprint arXiv:2511.03213},
  year   = {2025}
}

Comments

Accepted by CSF 2026 -- 39th IEEE Computer Security Foundations Symposium

R2 v1 2026-07-01T07:22:26.309Z