English
Related papers

Related papers: Neural Text Sanitization with Privacy Risk Indicat…

200 papers

In this paper we apply self-knowledge distillation to text summarization which we argue can alleviate problems with maximum-likelihood training on single reference and noisy datasets. Instead of relying on one-hot annotation labels, our…

Computation and Language · Computer Science 2021-07-28 Yang Liu , Sheng Shen , Mirella Lapata

There is an increasing concern in computer vision devices invading users' privacy by recording unwanted videos. On the one hand, we want the camera systems to recognize important events and assist human daily lives by understanding its…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Zhongzheng Ren , Yong Jae Lee , Michael S. Ryoo

Redactable signature schemes and sanitizable signature schemes are methods that permit modification of a given digital message and retain a valid signature. This can be applied to decentralized identity systems for delegating identity…

Cryptography and Security · Computer Science 2023-10-27 Bryan Kumara , Mark Hooper , Carsten Maple , Timothy Hobson , Jon Crowcroft

Speaker anonymization aims to suppress speaker individuality to protect privacy in speech while preserving the other aspects, such as speech content. One effective solution for anonymization is to modify the McAdams coefficient. In this…

Cryptography and Security · Computer Science 2021-07-16 Candy Olivia Mawalim , Masashi Unoki

Current online translation services require sending user text to cloud servers, posing a risk of privacy leakage when the text contains sensitive information. This risk hinders the application of online translation services in…

Computation and Language · Computer Science 2026-03-17 Wei Shao , Lemao Liu , Yinqiao Li , Guoping Huang , Shuming Shi , Linqi Song

Operators of online social networks are increasingly sharing potentially sensitive information about users and their relationships with advertisers, application developers, and data-mining researchers. Privacy is typically protected by…

Cryptography and Security · Computer Science 2016-11-17 Arvind Narayanan , Vitaly Shmatikov

Most tasks in NLP require labeled data. Data labeling is often done on crowdsourcing platforms due to scalability reasons. However, publishing data on public platforms can only be done if no privacy-relevant information is included. Textual…

Computation and Language · Computer Science 2023-03-07 Nina Mouhammad , Johannes Daxenberger , Benjamin Schiller , Ivan Habernal

Web search logs contain extremely sensitive data, as evidenced by the recent AOL incident. However, storing and analyzing search logs can be very useful for many purposes (i.e. investigating human behavior). Thus, an important research…

Databases · Computer Science 2015-03-19 Yuan Hong , Jaideep Vaidya , Haibing Lu , Mingrui Wu

Text rewriting with differential privacy (DP) provides concrete theoretical guarantees for protecting the privacy of individuals in textual documents. In practice, existing systems may lack the means to validate their privacy-preserving…

Computation and Language · Computer Science 2022-08-23 Timour Igamberdiev , Thomas Arnold , Ivan Habernal

With the popularity of virtual assistants (e.g., Siri, Alexa), the use of speech recognition is now becoming more and more widespread.However, speech signals contain a lot of sensitive information, such as the speaker's identity, which…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-21 Pierre Champion , Denis Jouvet , Anthony Larcher

Detecting personal health mentions on social media is essential to complement existing health surveillance systems. However, annotating data for detecting health mentions at a large scale is a challenging task. This research employs a…

Computation and Language · Computer Science 2022-12-13 Olanrewaju Tahir Aduragba , Jialin Yu , Alexandra I. Cristea

Text segmentation, the task of dividing a document into contiguous segments based on its semantic structure, is a longstanding challenge in language understanding. Previous work on text segmentation focused on unsupervised methods such as…

Computation and Language · Computer Science 2018-03-28 Omri Koshorek , Adir Cohen , Noam Mor , Michael Rotman , Jonathan Berant

This work addresses critical challenges to academic integrity, including plagiarism, fabrication, and verification of authorship of educational content, by proposing a Natural Language Processing (NLP)-based framework for authenticating…

To enable process analysis based on an event log without compromising the privacy of individuals involved in process execution, a log may be anonymized. Such anonymization strives to transform a log so that it satisfies provable privacy…

Cryptography and Security · Computer Science 2021-08-11 Fabian Rösel , Stephan A. Fahrenkrog-Petersen , Han van der Aa , Matthias Weidlich

Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has…

Computation and Language · Computer Science 2025-07-11 Dominykas Seputis , Yongkang Li , Karsten Langerak , Serghei Mihailov

Publishing person-specific transactions in an anonymous form is increasingly required by organizations. Recent approaches ensure that potentially identifying information (e.g., a set of diagnosis codes) cannot be used to link published…

Databases · Computer Science 2010-01-26 Grigorios Loukides , Aris Gkoulalas-Divanis , Bradley Malin

Recent literature has seen a considerable uptick in $\textit{Differentially Private Natural Language Processing}$ (DP NLP). This includes DP text privatization, where potentially sensitive input texts are transformed under DP to achieve…

Computation and Language · Computer Science 2025-03-13 Stephen Meisenbacher , Alexandra Klymenko , Alexander Karpp , Florian Matthes

Over the recent years, the availability of datasets containing personal, but anonymized information has been continuously increasing. Extensive research has revealed that such datasets are vulnerable to privacy breaches: being able to…

Cryptography and Security · Computer Science 2019-02-27 Alexandros Bampoulidis , Mihai Lupu

Recent work has demonstrated the successful extraction of training data from generative language models. However, it is not evident whether such extraction is feasible in text classification models since the training objective is to predict…

Computation and Language · Computer Science 2022-06-10 Adel Elmahdy , Huseyin A. Inan , Robert Sim

De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data. It has been well-studied within the medical domain. The need for de-identification technology is increasing, as…

Computation and Language · Computer Science 2021-05-25 Kristian Nørgaard Jensen , Mike Zhang , Barbara Plank