中文

面向 Topics API 输出的差分隐私合成数据释放

密码学与安全 2025-07-01 v1 人工智能 机器学习

摘要

隐私保留广告 API 隐私属性的分析是已获得学术界、行业界和监管机构高度关注的领域。尽管如此,对这些方法的实证研究仍受限于缺乏公开可用数据。可靠地对 API 的隐私属性进行实证分析,实际需要获得由真实 API 输出组成的数据集;然而,隐私顾虑阻碍了此类数据向公众的普遍释放。在本工作中,我们发展了一种新方法论,用于构建同时具备足够真实性以 enable accurate study 且提供强隐私保护的合成 API 输出。我们聚焦于一种隐私保留广告 API:Google Chrome 隐私沙箱的一部分 Topics API。我们的方法论基于首先计算大量描述 API 轨迹随时间演变的差分隐私统计量。随后,我们设计一个关于 API 轨迹序列的可参数化分布,并优化其参数,以使其与获取的统计量高度匹配。最后,我们通过从该分布中抽取来创建合成数据。 our work is complemented by an open-source release of the anonymized dataset obtained by this methodology. 我们希望这将使外部研究者能够深入分析该 API 并在真实大规模数据集上复制以前及未来工作。我们相信本工作将促进对隐私保留广告 API 隐私属性的透明度。

关键词

引用

@article{arxiv.2506.23855,
  title  = {Differentially Private Synthetic Data Release for Topics API Outputs},
  author = {Travis Dick and Alessandro Epasto and Adel Javanmard and Josh Karlin and Andres Munoz Medina and Vahab Mirrokni and Sergei Vassilvitskii and Peilin Zhong},
  journal= {arXiv preprint arXiv:2506.23855},
  year   = {2025}
}

备注

20 pages, 8 figures