中文
相关论文

相关论文: Textwash -- automated open-source text anonymisati…

200 篇论文

The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered…

计算与语言 · 计算机科学 2026-05-26 Angel Paul , Dhivin Shaji , Lifeng Han , Warren Del-Pinto , Goran Nenadic , Suzan Verberne

Named Entity Recognition and Disambiguation (NERD) systems are foundational for information retrieval, question answering, event detection, and other natural language processing (NLP) applications. We introduce TweetNERD, a dataset of 340K+…

计算与语言 · 计算机科学 2022-10-18 Shubhanshu Mishra , Aman Saini , Raheleh Makki , Sneha Mehta , Aria Haghighi , Ali Mollahosseini

Social networks have become an essential meeting point for millions of individuals willing to publish and consume huge quantities of heterogeneous information. Some studies have shown that the data published in these platforms may contain…

密码学与安全 · 计算机科学 2016-07-05 Alexandre Viejo , David Sánchez

Responsible use of AI demands that we protect sensitive information without undermining the usefulness of data, an imperative that has become acute in the age of large language models. We address this challenge with an on-premise,…

计算与语言 · 计算机科学 2026-03-19 Federico Albanese , Pablo Ronco , Nicolás D'Ippolito

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number…

We present a free and open-source tool for creating web-based surveys that include text annotation tasks. Existing tools offer either text annotation or survey functionality but not both. Combining the two input types is particularly…

计算与语言 · 计算机科学 2021-12-20 Timo Spinde , Kanishka Sinha , Norman Meuschke , Bela Gipp

As mobile app usage continues to rise, so does the generation of extensive user interaction data, which includes actions such as swiping, zooming, or the time spent on a screen. Apps often collect a large amount of this data and claim to…

软件工程 · 计算机科学 2024-04-24 Feiyang Tang , Bjarte M. Østvold

Clinical text processing has gained more and more attention in recent years. The access to sensitive patient data, on the other hand, is still a big challenge, as text cannot be shared without legal hurdles and without removing personal…

计算与语言 · 计算机科学 2022-09-02 Iyadh Ben Cheikh Larbi , Aljoscha Burchardt , Roland Roller

Our ability to control the flow of sensitive personal information to online systems is key to trust in personal privacy on the internet. We ask how to detect, assess and defend user privacy in the face of search engine personalisation? We…

密码学与安全 · 计算机科学 2016-09-27 Pól Mac Aonghusa , Douglas J. Leith

With Internet users constantly leaving a trail of text, whether through blogs, emails, or social media posts, the ability to write and protest anonymously is being eroded because artificial intelligence, when given a sample of previous…

机器学习 · 计算机科学 2021-10-19 Rishi Balakrishnan , Stephen Sloan , Anil Aswani

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or replacing sensitive information from data, enabling its…

计算与语言 · 计算机科学 2020-03-18 Aitor García-Pablos , Naiara Perez , Montse Cuadros

The increasing use of cloud-based speech assistants has heightened the need for effective speech anonymization, which aims to obscure a speaker's identity while retaining critical information for subsequent tasks. One approach to achieving…

人工智能 · 计算机科学 2024-10-22 Suhita Ghosh , Tim Thiele , Frederic Lorbeer , Frank Dreyer , Sebastian Stober

Anonymized data is highly valuable to both businesses and researchers. A large body of research has however shown the strong limits of the de-identification release-and-forget model, where data is anonymized and shared. This has led to the…

密码学与安全 · 计算机科学 2019-10-31 Andrea Gadotti , Florimond Houssiau , Luc Rocher , Benjamin Livshits , Yves-Alexandre de Montjoye

To enable process analysis based on an event log without compromising the privacy of individuals involved in process execution, a log may be anonymized. Such anonymization strives to transform a log so that it satisfies provable privacy…

密码学与安全 · 计算机科学 2021-08-11 Fabian Rösel , Stephan A. Fahrenkrog-Petersen , Han van der Aa , Matthias Weidlich

The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in…

计算与语言 · 计算机科学 2023-03-24 Xavier Tannier , Perceval Wajsbürt , Alice Calliger , Basile Dura , Alexandre Mouchet , Martin Hilka , Romain Bey

In this paper we propose use of a k-anonymity-like approach for evaluating the privacy of redacted text. Given a piece of redacted text we use a state of the art transformer-based deep learning network to reconstruct the original text. This…

机器学习 · 计算机科学 2024-10-11 Vaibhav Gusain , Douglas Leith

Social media plays an important role for a vast majority in one's internet life. Likewise, sharing, publishing and posting content through social media became nearly effortless. This unleashes new threats as unintentionally shared…

信息检索 · 计算机科学 2022-08-19 Stefan Kutschera

Social scientists are increasingly interested in analyzing the semantic information (e.g., emotion) of unstructured data (e.g., Tweets), where the semantic information is not natively present. Performing this analysis in a cost-efficient…

数据库 · 计算机科学 2025-01-08 Chuxuan Hu , Austin Peters , Daniel Kang

Voice data generated on instant messaging or social media applications contains unique user voiceprints that may be abused by malicious adversaries for identity inference or identity theft. Existing voice anonymization techniques, e.g.,…

声音 · 计算机科学 2022-10-28 Jiangyi Deng , Fei Teng , Yanjiao Chen , Xiaofu Chen , Zhaohui Wang , Wenyuan Xu

Predicting personality is essential for social applications supporting human-centered activities, yet prior modeling methods with users written text require too much input data to be realistically used in the context of social media. In…

社会与信息网络 · 计算机科学 2017-04-20 Pierre-Hadrien Arnoux , Anbang Xu , Neil Boyette , Jalal Mahmud , Rama Akkiraju , Vibha Sinha