中文
相关论文

相关论文: Textwash -- automated open-source text anonymisati…

200 篇论文

Anonymizing text that contains sensitive information is crucial for a wide range of applications. Existing techniques face the emerging challenges of the re-identification ability of large language models (LLMs), which have shown advanced…

计算与语言 · 计算机科学 2025-06-19 Tianyu Yang , Xiaodan Zhu , Iryna Gurevych

In today's digital world, casual user-generated content often contains subtle cues that may inadvertently expose sensitive personal attributes. Such risks underscore the growing importance of effective text anonymization to safeguard…

计算与语言 · 计算机科学 2025-07-01 Chenyang Shao , Tianxing Li , Chenhao Pu , Fengli Xu , Yong Li

Open science is a fundamental pillar to promote scientific progress and collaboration, based on the principles of open data, open source and open access. However, the requirements for publishing and sharing open data are in many cases…

密码学与安全 · 计算机科学 2024-08-21 Judith Sáinz-Pardo Díaz , Álvaro López García

Anonymization plays a key role in protecting sensible information of individuals in real world datasets. Self-driving cars for example need high resolution facial features to track people and their viewing direction to predict future…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Pascal Zwick , Kevin Roesch , Marvin Klemp , Oliver Bringmann

Texts convey sophisticated knowledge. However, texts also convey sensitive information. Despite the success of general-purpose language models and domain-specific mechanisms with differential privacy (DP), existing text sanitization…

计算与语言 · 计算机科学 2021-06-03 Xiang Yue , Minxin Du , Tianhao Wang , Yaliang Li , Huan Sun , Sherman S. M. Chow

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

计算与语言 · 计算机科学 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

Clinical free-text data offers immense potential to improve population health research such as richer phenotyping, symptom tracking, and contextual understanding of patient care. However, these data present significant privacy risks due to…

We propose a novel method to bootstrap text anonymization models based on distant supervision. Instead of requiring manually labeled training data, the approach relies on a knowledge graph expressing the background information assumed to be…

计算与语言 · 计算机科学 2022-05-17 Anthi Papadopoulou , Pierre Lison , Lilja Øvrelid , Ildikó Pilán

Over the recent years, the availability of datasets containing personal, but anonymized information has been continuously increasing. Extensive research has revealed that such datasets are vulnerable to privacy breaches: being able to…

密码学与安全 · 计算机科学 2019-02-27 Alexandros Bampoulidis , Mihai Lupu

Operators of online social networks are increasingly sharing potentially sensitive information about users and their relationships with advertisers, application developers, and data-mining researchers. Privacy is typically protected by…

密码学与安全 · 计算机科学 2016-11-17 Arvind Narayanan , Vitaly Shmatikov

We show that while anonymization effectively obscures firm identity, it significantly reduces the power of textual understanding, thereby diminishing models' ability to extract meaningful economic signals from financial texts. This…

综合金融 · 定量金融 2025-11-20 Ke Wu , Baozhong Yang , Zhenkun Ying , Dexin Zhou

Anonymizing sensitive information in user text is essential for privacy, yet existing methods often apply uniform treatment across attributes, which can conflict with communicative intent and obscure necessary information. This is…

密码学与安全 · 计算机科学 2026-01-09 Weihao Shen , Yaxin Xu , Shuang Li , Wei Chen , Yuqin Lan , Meng Yuan , Fuzhen Zhuang

Machine Learning approaches to Natural Language Processing tasks benefit from a comprehensive collection of real-life user data. At the same time, there is a clear need for protecting the privacy of the users whose data is collected and…

计算与语言 · 计算机科学 2022-11-16 David Ifeoluwa Adelani , Ali Davody , Thomas Kleinbauer , Dietrich Klakow

Unstructured text from legal, medical, and administrative sources offers a rich but underutilized resource for research in public health and the social sciences. However, large-scale analysis is hampered by two key challenges: the presence…

计算与语言 · 计算机科学 2025-07-16 Anders Ledberg , Anna Thalén

Large Language Models (LLMs) have demonstrated advanced capabilities in both text generation and comprehension, and their application to data archives might facilitate the privatization of sensitive information about the data subjects. In…

密码学与安全 · 计算机科学 2025-04-08 Stefano Cirillo , Domenico Desiato , Giuseppe Polese , Monica Maria Lucia Sebillo , Giandomenico Solimando

An unsolved challenge in distributed or federated learning is to effectively mitigate privacy risks without slowing down training or reducing accuracy. In this paper, we propose TextHide aiming at addressing this challenge for natural…

计算与语言 · 计算机科学 2020-10-14 Yangsibo Huang , Zhao Song , Danqi Chen , Kai Li , Sanjeev Arora

Recent advances in text mining and natural language processing technology have enabled researchers to detect an authors identity or demographic characteristics, such as age and gender, in several text genres by automatically analysing the…

密码学与安全 · 计算机科学 2022-11-30 Claudia Peersman , Matthew Edwards , Emma Williams , Awais Rashid

Authorship obfuscation techniques hold the promise of helping people protect their privacy in online communications by automatically rewriting text to hide the identity of the original author. However, obfuscation has been evaluated in…

计算与语言 · 计算机科学 2024-05-17 Calvin Bao , Marine Carpuat

We explore the feasibility of automatically finding accounts that publish sensitive content on Twitter. One natural approach to this problem is to first create a list of sensitive keywords, and then identify Twitter accounts that use these…

社会与信息网络 · 计算机科学 2017-02-02 Sai Teja Peddinti , Keith W. Ross , Justin Cappos

Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still catergorized as non-sensitive. Indeed, in the past, researchers and…

密码学与安全 · 计算机科学 2018-05-15 Beatrice Perez , Mirco Musolesi , Gianluca Stringhini