中文
相关论文

相关论文: Textwash -- automated open-source text anonymisati…

200 篇论文

Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds…

计算与语言 · 计算机科学 2025-05-23 Matthew Zent , Digory Smith , Simon Woodhead

Synthetic data is increasingly used to support research without exposing sensitive user content. Social media data is one of the types of datasets that would hugely benefit from representative synthetic equivalents that can be used to…

密码学与安全 · 计算机科学 2026-03-06 Henry Tari , Adriana Iamnitchi

Data anonymization is an approach to privacy-preserving data release aimed at preventing participants reidentification, and it is an important alternative to differential privacy in applications that cannot tolerate noisy data. Existing…

数据结构与算法 · 计算机科学 2022-01-31 Gecia Bravo-Hermsdorff , Robert Busa-Fekete , Lee M. Gunderson , Andrés Munõz Medina , Umar Syed

Anonymization of graph-based data is a problem which has been widely studied over the last years and several anonymization methods have been developed. Information loss measures have been used to evaluate data utility and information loss…

密码学与安全 · 计算机科学 2025-02-03 Jordi Casas-Roma

Exploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept…

计算与语言 · 计算机科学 2020-05-20 Lukas Lange , Heike Adel , Jannik Strötgen

In medical organizations large amount of personal data are collected and analyzed by the data miner or researcher, for further perusal. However, the data collected may contain sensitive information such as specific disease of a patient and…

密码学与安全 · 计算机科学 2012-03-19 Pawan R Bhaladhare , Devesh Jinwala

To be informative, an evaluation must measure how well systems generalize to realistic unseen data. We identify limitations of and propose improvements to current evaluations of text-to-SQL systems. First, we compare human-generated and…

Today, the publication of microdata poses a privacy threat. Vast research has striven to define the privacy condition that microdata should satisfy before it is released, and devise algorithms to anonymize the data so as to achieve this…

数据库 · 计算机科学 2012-08-02 Jianneng Cao , Panagiotis Karras

With the increasing use of social media data for health-related research, the credibility of the information from this source has been questioned as the posts may originate from automated accounts or "bots". While automatic bot detection…

计算与语言 · 计算机科学 2019-10-01 Anahita Davoudi , Ari Z. Klein , Abeed Sarker , Graciela Gonzalez-Hernandez

As large language models (LLMs) rapidly advance and integrate into daily life, the privacy risks they pose are attracting increasing attention. We focus on a specific privacy risk where LLMs may help identify the authorship of anonymous…

计算与语言 · 计算机科学 2024-11-21 Zichen Wen , Dadi Guo , Huishuai Zhang

Speaker anonymization aims to conceal a speaker's identity, without considering the linguistic content. In this study, we reveal a weakness of Librispeech, the dataset that is commonly used to evaluate anonymizers: the books read by the…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Carlos Franzreb , Arnab Das , Tim Polzehl , Sebastian Möller

This paper considers the problem of automatically characterizing overall attitudes and biases that may be associated with emerging information operations via artificial intelligence. Accurate analysis of these emerging topics usually…

计算机与社会 · 计算机科学 2020-12-07 Autumn Toney , Akshat Pandey , Wei Guo , David Broniatowski , Aylin Caliskan

The risks of publishing privacy-sensitive data have received considerable attention recently. Several de-anonymization attacks have been proposed to re-identify individuals even if data anonymization techniques were applied. However, there…

社会与信息网络 · 计算机科学 2017-03-16 Wei-Han Lee , Changchang Liu , Shouling Ji , Prateek Mittal , Ruby Lee

Social media platforms, increasingly used as news sources for varied data analytics, have transformed how information is generated and disseminated. However, the unverified nature of this content raises concerns about trustworthiness and…

Linguistic bias in online news and social media is widespread but difficult to measure. Yet, its identification and quantification remain difficult due to subjectivity, context dependence, and the scarcity of high-quality gold-label…

信息检索 · 计算机科学 2025-12-17 Fabian Haak , Philipp Schaer

It is well known that textual data on the internet and other digital platforms contain significant levels of bias and stereotypes. Although many such texts contain stereotypes and biases that inherently exist in natural language for reasons…

计算与语言 · 计算机科学 2022-01-24 Ewoenam Kwaku Tokpo , Toon Calders

Faced with the threat of identity leakage during voice data publishing, users are engaged in a privacy-utility dilemma when enjoying convenient voice services. Existing studies employ direct modification or text-based re-synthesis to…

声音 · 计算机科学 2022-11-11 Meng Chen , Li Lu , Jiadi Yu , Yingying Chen , Zhongjie Ba , Feng Lin , Kui Ren

The rapid advancement of large language models (LLMs) has enabled powerful authorship inference capabilities, raising growing concerns about unintended deanonymization risks in textual data such as news articles. In this work, we introduce…

计算与语言 · 计算机科学 2026-02-27 Boyang Zhang , Yang Zhang

Large Language Models (LLMs) are gaining increasing attention due to their exceptional performance across numerous tasks. As a result, the general public utilize them as an influential tool for boosting their productivity while natural…

密码学与安全 · 计算机科学 2023-06-16 Zhigang Kan , Linbo Qiao , Hao Yu , Liwen Peng , Yifu Gao , Dongsheng Li

Sharing real-world speech utterances is key to the training and deployment of voice-based services. However, it also raises privacy risks as speech contains a wealth of personal data. Speaker anonymization aims to remove speaker information…