中文
相关论文

相关论文: Neural Text Sanitization with Privacy Risk Indicat…

200 篇论文

In many countries, personal information that can be published or shared between organizations is regulated and, therefore, documents must undergo a process of de-identification to eliminate or obfuscate confidential data. Our work focuses…

计算与语言 · 计算机科学 2019-10-10 Diego Garat , Dina Wonsever

Social networks have become an essential meeting point for millions of individuals willing to publish and consume huge quantities of heterogeneous information. Some studies have shown that the data published in these platforms may contain…

密码学与安全 · 计算机科学 2016-07-05 Alexandre Viejo , David Sánchez

Anonymization is a foundational principle of data privacy regulation, yet its practical application remains riddled with ambiguity and inconsistency. This paper introduces the concept of anonymity-washing -- the misrepresentation of the…

密码学与安全 · 计算机科学 2025-08-27 Szivia Lestyán , William Letrone , Ludovica Robustelli , Gergely Biczók

Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable information (PII), thereby making it more difficult to…

计算与语言 · 计算机科学 2025-05-20 Lucas Georges Gabriel Charpentier , Pierre Lison

The literature on data sanitization aims to design algorithms that take an input dataset and produce a privacy-preserving version of it, that captures some of its statistical properties. In this note we study this question from a streaming…

数据结构与算法 · 计算机科学 2021-11-30 Haim Kaplan , Uri Stemmer

We propose a novel redaction methodology that can be used to sanitize natural text data. Our new technique provides better privacy benefits than other state of the art techniques while maintaining lower redaction levels.

密码学与安全 · 计算机科学 2025-06-19 Vaibhav Gusain , Douglas Leith

Most of privacy protection studies for textual data focus on removing explicit sensitive identifiers. However, personal writing style, as a strong indicator of the authorship, is often neglected. Recent studies, such as SynTF, have shown…

密码学与安全 · 计算机科学 2021-05-14 Haohan Bo , Steven H. H. Ding , Benjamin C. M. Fung , Farkhund Iqbal

The widespread use of cloud-based Large Language Models (LLMs) has heightened concerns over user privacy, as sensitive information may be inadvertently exposed during interactions with these services. To protect privacy before sending…

计算与语言 · 计算机科学 2025-05-28 Shuo Huang , William MacLean , Xiaoxi Kang , Qiongkai Xu , Zhuang Li , Xingliang Yuan , Gholamreza Haffari , Lizhen Qu

The extensive use of online social media has highlighted the importance of privacy in the digital space. As more scientists analyse the data created in these platforms, privacy concerns have extended to data usage within the academia.…

人机交互 · 计算机科学 2022-03-04 Giannis Haralabopoulos , Ioannis Anagnostopoulos

We address the problem of how to "obfuscate" texts by removing stylistic clues which can identify authorship, whilst preserving (as much as possible) the content of the text. In this paper we combine ideas from "generalised differential…

密码学与安全 · 计算机科学 2019-02-06 Natasha Fernandes , Mark Dras , Annabelle McIver

Differential Privacy (DP) for text matured from disjointed word-level substitutions to contiguous sentence-level rewriting by leveraging the generative capacity of language models. While this form of text privatization is best suited for…

计算与语言 · 计算机科学 2026-04-30 Stefan Arnold

Online users generate tremendous amounts of textual information by participating in different activities, such as writing reviews and sharing tweets. This textual data provides opportunities for researchers and business partners to study…

密码学与安全 · 计算机科学 2019-07-09 Ghazaleh Beigi , Kai Shu , Ruocheng Guo , Suhang Wang , Huan Liu

The steadily increasing utilization of data-driven methods and approaches in areas that handle sensitive personal information such as in law enforcement mandates an ever increasing effort in these institutions to comply with data protection…

人工智能 · 计算机科学 2025-01-14 Manuel Eberhardinger , Patrick Takenaka , Daniel Grießhaber , Johannes Maucher

The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and…

音频与语音处理 · 电气工程与系统科学 2020-09-01 Paul-Gauthier Noé , Jean-François Bonastre , Driss Matrouf , Natalia Tomashenko , Andreas Nautsch , Nicholas Evans

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

计算与语言 · 计算机科学 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

Adolescent suicide is a critical global health issue, and speech provides a cost-effective modality for automatic suicide risk detection. Given the vulnerable population, protecting speaker identity is particularly important, as speech…

音频与语音处理 · 电气工程与系统科学 2026-01-26 Ziyun Cui , Sike Jia , Yang Lin , Yinan Duan , Diyang Qu , Runsen Chen , Chao Zhang , Chang Lei , Wen Wu

Clinical text processing has gained more and more attention in recent years. The access to sensitive patient data, on the other hand, is still a big challenge, as text cannot be shared without legal hurdles and without removing personal…

计算与语言 · 计算机科学 2022-09-02 Iyadh Ben Cheikh Larbi , Aljoscha Burchardt , Roland Roller

We propose sanitizer, a framework for secure and task-agnostic data release. While releasing datasets continues to make a big impact in various applications of computer vision, its impact is mostly realized when data sharing is not…

密码学与安全 · 计算机科学 2022-03-25 Abhishek Singh , Ethan Garza , Ayush Chopra , Praneeth Vepakomma , Vivek Sharma , Ramesh Raskar

The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving…

机器学习 · 计算机科学 2025-07-08 Jing Yu , Yibo Zhao , Jiapeng Zhu , Wenming Shao , Bo Pang , Zhao Zhang , Xiang Li

The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation. Standard anonymization techniques often…

密码学与安全 · 计算机科学 2026-01-09 Lionel Z. Wang , Yusheng Zhao , Jiabin Luo , Xinfeng Li , Lixu Wang , Yinan Peng , Haoyang Li , XiaoFeng Wang , Wei Dong