中文
相关论文

相关论文: Neural Text Sanitization with Privacy Risk Indicat…

200 篇论文

In this paper we apply self-knowledge distillation to text summarization which we argue can alleviate problems with maximum-likelihood training on single reference and noisy datasets. Instead of relying on one-hot annotation labels, our…

计算与语言 · 计算机科学 2021-07-28 Yang Liu , Sheng Shen , Mirella Lapata

There is an increasing concern in computer vision devices invading users' privacy by recording unwanted videos. On the one hand, we want the camera systems to recognize important events and assist human daily lives by understanding its…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Zhongzheng Ren , Yong Jae Lee , Michael S. Ryoo

Redactable signature schemes and sanitizable signature schemes are methods that permit modification of a given digital message and retain a valid signature. This can be applied to decentralized identity systems for delegating identity…

密码学与安全 · 计算机科学 2023-10-27 Bryan Kumara , Mark Hooper , Carsten Maple , Timothy Hobson , Jon Crowcroft

Speaker anonymization aims to suppress speaker individuality to protect privacy in speech while preserving the other aspects, such as speech content. One effective solution for anonymization is to modify the McAdams coefficient. In this…

密码学与安全 · 计算机科学 2021-07-16 Candy Olivia Mawalim , Masashi Unoki

Current online translation services require sending user text to cloud servers, posing a risk of privacy leakage when the text contains sensitive information. This risk hinders the application of online translation services in…

计算与语言 · 计算机科学 2026-03-17 Wei Shao , Lemao Liu , Yinqiao Li , Guoping Huang , Shuming Shi , Linqi Song

Operators of online social networks are increasingly sharing potentially sensitive information about users and their relationships with advertisers, application developers, and data-mining researchers. Privacy is typically protected by…

密码学与安全 · 计算机科学 2016-11-17 Arvind Narayanan , Vitaly Shmatikov

Most tasks in NLP require labeled data. Data labeling is often done on crowdsourcing platforms due to scalability reasons. However, publishing data on public platforms can only be done if no privacy-relevant information is included. Textual…

计算与语言 · 计算机科学 2023-03-07 Nina Mouhammad , Johannes Daxenberger , Benjamin Schiller , Ivan Habernal

Web search logs contain extremely sensitive data, as evidenced by the recent AOL incident. However, storing and analyzing search logs can be very useful for many purposes (i.e. investigating human behavior). Thus, an important research…

数据库 · 计算机科学 2015-03-19 Yuan Hong , Jaideep Vaidya , Haibing Lu , Mingrui Wu

Text rewriting with differential privacy (DP) provides concrete theoretical guarantees for protecting the privacy of individuals in textual documents. In practice, existing systems may lack the means to validate their privacy-preserving…

计算与语言 · 计算机科学 2022-08-23 Timour Igamberdiev , Thomas Arnold , Ivan Habernal

With the popularity of virtual assistants (e.g., Siri, Alexa), the use of speech recognition is now becoming more and more widespread.However, speech signals contain a lot of sensitive information, such as the speaker's identity, which…

音频与语音处理 · 电气工程与系统科学 2022-03-21 Pierre Champion , Denis Jouvet , Anthony Larcher

Detecting personal health mentions on social media is essential to complement existing health surveillance systems. However, annotating data for detecting health mentions at a large scale is a challenging task. This research employs a…

计算与语言 · 计算机科学 2022-12-13 Olanrewaju Tahir Aduragba , Jialin Yu , Alexandra I. Cristea

Text segmentation, the task of dividing a document into contiguous segments based on its semantic structure, is a longstanding challenge in language understanding. Previous work on text segmentation focused on unsupervised methods such as…

计算与语言 · 计算机科学 2018-03-28 Omri Koshorek , Adir Cohen , Noam Mor , Michael Rotman , Jonathan Berant

This work addresses critical challenges to academic integrity, including plagiarism, fabrication, and verification of authorship of educational content, by proposing a Natural Language Processing (NLP)-based framework for authenticating…

To enable process analysis based on an event log without compromising the privacy of individuals involved in process execution, a log may be anonymized. Such anonymization strives to transform a log so that it satisfies provable privacy…

密码学与安全 · 计算机科学 2021-08-11 Fabian Rösel , Stephan A. Fahrenkrog-Petersen , Han van der Aa , Matthias Weidlich

Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has…

计算与语言 · 计算机科学 2025-07-11 Dominykas Seputis , Yongkang Li , Karsten Langerak , Serghei Mihailov

Publishing person-specific transactions in an anonymous form is increasingly required by organizations. Recent approaches ensure that potentially identifying information (e.g., a set of diagnosis codes) cannot be used to link published…

数据库 · 计算机科学 2010-01-26 Grigorios Loukides , Aris Gkoulalas-Divanis , Bradley Malin

Recent literature has seen a considerable uptick in $\textit{Differentially Private Natural Language Processing}$ (DP NLP). This includes DP text privatization, where potentially sensitive input texts are transformed under DP to achieve…

计算与语言 · 计算机科学 2025-03-13 Stephen Meisenbacher , Alexandra Klymenko , Alexander Karpp , Florian Matthes

Over the recent years, the availability of datasets containing personal, but anonymized information has been continuously increasing. Extensive research has revealed that such datasets are vulnerable to privacy breaches: being able to…

密码学与安全 · 计算机科学 2019-02-27 Alexandros Bampoulidis , Mihai Lupu

Recent work has demonstrated the successful extraction of training data from generative language models. However, it is not evident whether such extraction is feasible in text classification models since the training objective is to predict…

计算与语言 · 计算机科学 2022-06-10 Adel Elmahdy , Huseyin A. Inan , Robert Sim

De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data. It has been well-studied within the medical domain. The need for de-identification technology is increasing, as…

计算与语言 · 计算机科学 2021-05-25 Kristian Nørgaard Jensen , Mike Zhang , Barbara Plank