中文
相关论文

相关论文: On the Vulnerability of Text Sanitization

200 篇论文

Web tracking through third-party cookies is considered a threat to users' privacy and is supposed to be abandoned in the near future. Recently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural…

计算机与社会 · 计算机科学 2023-06-09 Nikhil Jha , Martino Trevisan , Emilio Leonardi , Marco Mellia

Preserving privacy of continuous and/or high-dimensional data such as images, videos and audios, can be challenging with syntactic anonymization methods which are designed for discrete attributes. Differential privacy, which provides a more…

机器学习 · 计算机科学 2017-12-04 Jihun Hamm

Authorship identification has proven unsettlingly effective in inferring the identity of the author of an unsigned document, even when sensitive personal information has been carefully omitted. In the digital era, individuals leave a…

计算与语言 · 计算机科学 2023-10-04 Haining Wang

We introduce a new method for extracting structured threat behaviors from threat intelligence text. Our method is based on a multi-stage ranking architecture that allows jointly optimizing for efficiency and effectiveness. Therefore, we…

密码学与安全 · 计算机科学 2024-03-27 Udesh Kumarasinghe , Ahmed Lekssays , Husrev Taha Sencar , Sabri Boughorbel , Charitha Elvitigala , Preslav Nakov

The rapid development of large language models (LLMs) has yielded impressive success in various downstream tasks. However, the vast potential and remarkable capabilities of LLMs also raise new security and privacy concerns if they are…

密码学与安全 · 计算机科学 2024-10-11 Jiawei Zhao , Kejiang Chen , Xiaojian Yuan , Yuang Qi , Weiming Zhang , Nenghai Yu

We propose reconstruction advantage measures to audit label privatization mechanisms. A reconstruction advantage measure quantifies the increase in an attacker's ability to infer the true label of an unlabeled example when provided with a…

Recent privacy research on large language models (LLMs) has shown that they achieve near-human-level performance at inferring personal data from online texts. With ever-increasing model capabilities, existing text anonymization methods are…

人工智能 · 计算机科学 2025-02-04 Robin Staab , Mark Vero , Mislav Balunović , Martin Vechev

This study investigates privacy leakage in dimensionality reduction methods through a novel machine learning-based reconstruction attack. Employing an informed adversary threat model, we develop a neural network capable of reconstructing…

密码学与安全 · 计算机科学 2025-06-03 Chayadon Lumbut , Donlapark Ponnoprat

Most of privacy protection studies for textual data focus on removing explicit sensitive identifiers. However, personal writing style, as a strong indicator of the authorship, is often neglected. Recent studies, such as SynTF, have shown…

密码学与安全 · 计算机科学 2021-05-14 Haohan Bo , Steven H. H. Ding , Benjamin C. M. Fung , Farkhund Iqbal

With the increasing importance of the internet in our day to day life, data security in web application has become very crucial. Ever increasing on line and real time transaction services have led to manifold rise in the problems associated…

数据库 · 计算机科学 2013-11-27 Vrushali Randhe , Archana Chougule , Debajyoti Mukhopadhyay

The use of third-party datasets and pre-trained machine learning models poses a threat to NLP systems due to possibility of hidden backdoor attacks. Existing attacks involve poisoning the data samples such as insertion of tokens or sentence…

计算与语言 · 计算机科学 2024-04-09 Irina Alekseevskaia , Konstantin Arkhipenko

The landscape of adversarial attacks against text classifiers continues to grow, with new attacks developed every year and many of them available in standard toolkits, such as TextAttack and OpenAttack. In response, there is a growing body…

The best practice to prevent Cross Site Scripting (XSS) attacks is to apply encoders to sanitize untrusted data. To balance security and functionality, encoders should be applied to match the web page context, such as HTML body, JavaScript,…

密码学与安全 · 计算机科学 2018-04-06 Mahmoud Mohammadi , Bei-Tseng Chu , Heather Richter Lipford

Predictions of certifiably robust classifiers remain constant in a neighborhood of a point, making them resilient to test-time attacks with a guarantee. In this work, we present a previously unrecognized threat to robust machine learning…

机器学习 · 计算机科学 2021-03-31 Akshay Mehra , Bhavya Kailkhura , Pin-Yu Chen , Jihun Hamm

The goal of differentially private text obfuscation is to obfuscate, or "perturb", input texts with Differential Privacy (DP) guarantees, such that the private output texts are quantifiably indistinguishable from the originals. While…

计算与语言 · 计算机科学 2026-05-05 Stephen Meisenbacher , Angelo Kleinert , Florian Matthes

Text-to-image diffusion models are pushing the boundaries of what generative AI can achieve in our lives. Beyond their ability to generate general images, new personalization techniques have been proposed to customize the pre-trained base…

计算机与社会 · 计算机科学 2024-10-15 Boheng Li , Yanhao Wei , Yankai Fu , Zhenting Wang , Yiming Li , Jie Zhang , Run Wang , Tianwei Zhang

The widespread use of cloud-based Large Language Models (LLMs) has heightened concerns over user privacy, as sensitive information may be inadvertently exposed during interactions with these services. To protect privacy before sending…

计算与语言 · 计算机科学 2025-05-28 Shuo Huang , William MacLean , Xiaoxi Kang , Qiongkai Xu , Zhuang Li , Xingliang Yuan , Gholamreza Haffari , Lizhen Qu

We address practical implementation of a risk-weighted pseudo posterior synthesizer for microdata dissemination with a new re-weighting strategy that maximizes utility of released synthetic data under at any level of formal privacy…

统计方法学 · 统计学 2022-05-02 Terrance D. Savitsky , Jingchen Hu , Matthew R. Williams

Watermarking is a key technique for detecting AI-generated text. In this work, we study its vulnerabilities and introduce the Smoothing Attack, a novel watermark removal method. By leveraging the relationship between the model's confidence…

机器学习 · 计算机科学 2025-02-06 Hongyan Chang , Hamed Hassani , Reza Shokri

Components of machine learning systems are not (yet) perceived as security hotspots. Secure coding practices, such as ensuring that no execution paths depend on confidential inputs, have not yet been adopted by ML developers. We initiate…

密码学与安全 · 计算机科学 2020-11-04 Zhen Sun , Roei Schuster , Vitaly Shmatikov