English
Related papers

Related papers: The Metadata Anonymization Toolkit

200 papers

Text sanitization is the task of redacting a document to mask all occurrences of (direct or indirect) personal identifiers, with the goal of concealing the identity of the individual(s) referred in it. In this paper, we consider a two-step…

Computation and Language · Computer Science 2023-10-24 Anthi Papadopoulou , Pierre Lison , Mark Anderson , Lilja Øvrelid , Ildikó Pilán

Anonymizing text that contains sensitive information is crucial for a wide range of applications. Existing techniques face the emerging challenges of the re-identification ability of large language models (LLMs), which have shown advanced…

Computation and Language · Computer Science 2025-06-19 Tianyu Yang , Xiaodan Zhu , Iryna Gurevych

The exploding rate of data publishing in our networked society has magnified the risk of sensitive information leakage and misuse, pushing the need to secure multimedia content from unintended exposure to potentially untrusted third…

Cryptography and Security · Computer Science 2025-09-16 Andrea Ciccotelli , Hanaa Abbas , Roberto Di Pietro

AI intensive systems that operate upon user data face the challenge of balancing data utility with privacy concerns. We propose the idea and present the prototype of an open-source tool called Privacy Utility Trade-off (PUT) Workbench which…

Cryptography and Security · Computer Science 2019-02-06 Saurabh Srivastava , Vinay P. Namboodiri , T. V. Prabhakar

Anonymizing textual documents is a highly context-sensitive problem: the appropriate balance between privacy protection and utility preservation varies with the data domain, privacy objectives, and downstream application. However, existing…

Computation and Language · Computer Science 2026-04-21 Gabriel Loiseau , Damien Sileo , Damien Riquet , Maxime Meyer , Marc Tommasi

We propose a novel problem formulation to address the privacy-utility tradeoff, specifically when dealing with two distinct user groups characterized by unique sets of private and utility attributes. Unlike previous studies that primarily…

Machine Learning · Computer Science 2024-09-12 Bishwas Mandal , George Amariucai , Shuangqing Wei

The Topics API for the web is Google's privacy-enhancing alternative to replace third-party cookies. Results of prior work have led to an ongoing discussion between Google and research communities about the capability of Topics to trade off…

Cryptography and Security · Computer Science 2024-08-16 Yohan Beugin , Patrick McDaniel

The proliferation of textual data containing sensitive personal information across various domains requires robust anonymization techniques to protect privacy and comply with regulations, while preserving data usability for diverse and…

Computation and Language · Computer Science 2025-12-17 Tobias Deußer , Lorenz Sparrenberg , Armin Berger , Max Hahnbück , Christian Bauckhage , Rafet Sifa

An anonymization technique for databases is proposed that employs Principal Component Analysis. The technique aims at releasing the least possible amount of information, while preserving the utility of the data released in response to…

Cryptography and Security · Computer Science 2019-03-29 Giuseppe D'Acquisto , Maurizio Naldi

The Privacy Sandbox, launched in 2019, is a series of proposals from Google to reduce ``cross-site and cross-app tracking while helping to keep online content and services free for all''. Over the years, Google implemented, experimented,…

Cryptography and Security · Computer Science 2025-12-04 Yohan Beugin , Patrick McDaniel

Within the current context of Information Societies, large amounts of information are daily exchanged and/or released. The sensitive nature of much of this information causes a serious privacy threat when documents are uncontrollably made…

Cryptography and Security · Computer Science 2017-07-07 David Sanchez , Montserrat Batet

Machine learning (ML) algorithms are heavily based on the availability of training data, which, depending on the domain, often includes sensitive information about data providers. This raises critical privacy concerns. Anonymization…

Machine Learning · Computer Science 2025-11-03 Héber H. Arcolezi , Mina Alishahi , Adda-Akram Bendoukha , Nesrine Kaaniche

Several anonymization techniques, such as generalization and bucketization, have been designed for privacy preserving microdata publishing. Recent work has shown that generalization loses considerable amount of information, especially for…

Databases · Computer Science 2009-09-15 Tiancheng Li , Ninghui Li , Jian Zhang , Ian Molloy

In medical organizations large amount of personal data are collected and analyzed by the data miner or researcher, for further perusal. However, the data collected may contain sensitive information such as specific disease of a patient and…

Cryptography and Security · Computer Science 2012-03-19 Pawan R Bhaladhare , Devesh Jinwala

Recently introduced privacy legislation has aimed to restrict and control the amount of personal data published by companies and shared to third parties. Much of this real data is not only sensitive requiring anonymization, but also…

Databases · Computer Science 2020-07-20 Mostafa Milani , Yu Huang , Fei Chiang

There are currently two approaches to anonymization: "utility first" (use an anonymization method with suitable utility features, then empirically evaluate the disclosure risk and, if necessary, reduce the risk by possibly sacrificing some…

Databases · Computer Science 2015-01-20 Josep Domingo-Ferrer , Krishnamurty Muralidhar

In a Multi-Agent System (MAS), individual agents observe various aspects of the environment and transmit this information to a central entity responsible for aggregating the data and deducing system parameters. To improve overall…

Cryptography and Security · Computer Science 2025-11-17 Puspanjali Ghoshal , Ashok Singh Sairam

Sanitizing sensitive text data typically involves removing personally identifiable information (PII) or generating synthetic data under the assumption that these methods adequately protect privacy; however, their effectiveness is often only…

Cryptography and Security · Computer Science 2026-03-17 Rui Xin , Niloofar Mireshghallah , Shuyue Stella Li , Michael Duan , Hyunwoo Kim , Yejin Choi , Yulia Tsvetkov , Sewoong Oh , Pang Wei Koh

Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to use, often with tools such as Microsoft's Presidio or…

Computation and Language · Computer Science 2026-02-16 Nataša Krčo , Zexi Yao , Matthieu Meeus , Yves-Alexandre de Montjoye

Enormous amounts of data collected from social networks or other online platforms are being published for the sake of statistics, marketing, and research, among other objectives. The consequent privacy and data security concerns have…

Cryptography and Security · Computer Science 2021-12-24 Ola N. Halawi , Faisal N. Abu-Khzam