English
Related papers

Related papers: Textwash -- automated open-source text anonymisati…

200 papers

Social media texts are significant information sources for several application areas including trend analysis, event monitoring, and opinion mining. Unfortunately, existing solutions for tasks such as named entity recognition that perform…

Computation and Language · Computer Science 2014-11-03 Dilek Küçük , Ralf Steinberger

This work investigates the effectiveness of different pseudonymization techniques, ranging from rule-based substitutions to using pre-trained Large Language Models (LLMs), on a variety of datasets and models used for two widely used NLP…

Computation and Language · Computer Science 2023-06-12 Oleksandr Yermilov , Vipul Raheja , Artem Chernodub

While social networks can provide an ideal platform for up-to-date information from individuals across the world, it has also proved to be a place where rumours fester and accidental or deliberate misinformation often emerges. In this…

Social and Information Networks · Computer Science 2016-11-22 Georgios Giasemidis , Colin Singleton , Ioannis Agrafiotis , Jason R. C. Nurse , Alan Pilgrim , Chris Willis , Danica Vukadinovic Greetham

Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes InfoLossQA, a framework to characterize and recover simplification-induced…

Computation and Language · Computer Science 2024-06-05 Jan Trienes , Sebastian Joseph , Jörg Schlötterer , Christin Seifert , Kyle Lo , Wei Xu , Byron C. Wallace , Junyi Jessy Li

In the realm of data privacy, the ability to effectively anonymise text is paramount. With the proliferation of deep learning and, in particular, transformer architectures, there is a burgeoning interest in leveraging these advanced models…

Doxing refers to the practice of disclosing sensitive personal information about a person without their consent. This form of cyberbullying is an unpleasant and sometimes dangerous phenomenon for online social networks. Although prior work…

Social and Information Networks · Computer Science 2022-11-15 Younes Karimi , Anna Squicciarini , Shomir Wilson

The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-01 Paul-Gauthier Noé , Jean-François Bonastre , Driss Matrouf , Natalia Tomashenko , Andreas Nautsch , Nicholas Evans

This work introduces an anonymization scheme for a corpus of texts to safeguard metadata from disclosure. It specifically aims to prevent large language models from identifying metadata associated with texts, thereby avoiding their…

Applications · Statistics 2025-05-28 Jan Greve , Lukas Sablica

The increased popularity and ubiquitous availability of online social networks and globalised Internet access have affected the way in which people share content. The information that users willingly disclose on these platforms can be used…

Social and Information Networks · Computer Science 2016-07-12 Maria Han Veiga , Carsten Eickhoff

Privacy analysis is critical but also a time-consuming and tedious task. We present a formalization which eases designing and auditing high-level privacy properties of software architectures. It is incorporated into a larger policy analysis…

Cryptography and Security · Computer Science 2018-06-11 Marcel von Maltitz , Cornelius Diekmann , Georg Carle

In many countries, personal information that can be published or shared between organizations is regulated and, therefore, documents must undergo a process of de-identification to eliminate or obfuscate confidential data. Our work focuses…

Computation and Language · Computer Science 2019-10-10 Diego Garat , Dina Wonsever

In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers. However, the absence of user identifiers restricts the…

Machine Learning · Computer Science 2025-01-13 Yucheng Ding , Yangwenjian Tan , Xiangyu Liu , Chaoyue Niu , Fandong Meng , Jie Zhou , Ning Liu , Fan Wu , Guihai Chen

Machine learning models can perpetuate unintended biases from unfair and imbalanced datasets. Evaluating and debiasing these datasets and models is especially hard in text datasets where sensitive attributes such as race, gender, and sexual…

Computation and Language · Computer Science 2024-01-15 Emmanuel Klu , Sameer Sethi

The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving…

Machine Learning · Computer Science 2025-07-08 Jing Yu , Yibo Zhao , Jiapeng Zhu , Wenming Shao , Bo Pang , Zhao Zhang , Xiang Li

In social media networks, users produce a large amount of text content anytime, providing researchers with an invaluable approach to digging for personality-related information. Personality detection based on user-generated text is a method…

Computers and Society · Computer Science 2025-09-18 Lei Lin , Jizhao Zhu , Qirui Tang , Yihua Du

Users posting online expect to remain anonymous unless they have logged in, which is often needed for them to be able to discuss freely on various topics. Preserving the anonymity of a text's writer can be also important in some other…

Computation and Language · Computer Science 2017-07-31 Georgi Karadjov , Tsvetomila Mihaylova , Yasen Kiprov , Georgi Georgiev , Ivan Koychev , Preslav Nakov

In this paper, de-anonymizing internet users by actively querying their group memberships in social networks is considered. In this problem, an anonymous victim visits the attacker's website, and the attacker uses the victim's browser…

Social and Information Networks · Computer Science 2018-01-22 F. Shirani , S. Garg , E. Erkip

Speaker anonymization is the task of modifying a speech recording such that the original speaker cannot be identified anymore. Since the first Voice Privacy Challenge in 2020, along with the release of a framework, the popularity of this…

Sound · Computer Science 2023-12-25 Sarina Meyer , Xiaoxiao Miao , Ngoc Thang Vu

Data sanitization in the context of language modeling involves identifying sensitive content, such as personally identifiable information (PII), and redacting them from a dataset corpus. It is a common practice used in natural language…

Computation and Language · Computer Science 2024-11-12 Anwesan Pal , Radhika Bhargava , Kyle Hinsz , Jacques Esterhuizen , Sudipta Bhattacharya

Although the bulk of the research in privacy and statistical disclosure control is designed for static data, more and more data are often collected as continuous streams, and extensions of popular privacy tools and models have been proposed…

Cryptography and Security · Computer Science 2024-02-27 Nicolas Ruiz