中文
相关论文

相关论文: The Text Anonymization Benchmark (TAB): A Dedicate…

200 篇论文

Biometric data is pervasively captured and analyzed. Using modern machine learning approaches, identity and attribute inferences attacks have proven high accuracy. Anonymizations aim to mitigate such disclosures by modifying data in a way…

密码学与安全 · 计算机科学 2024-07-10 Julian Todt , Simon Hanisch , Thorsten Strufe

Users today expect more security from services that handle their data. In addition to traditional data privacy and integrity requirements, they expect transparency, i.e., that the service's processing of the data is verifiable by users and…

密码学与安全 · 计算机科学 2022-10-24 Daniel Reijsbergen , Aung Maw , Zheng Yang , Tien Tuan Anh Dinh , Jianying Zhou

Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and complex, as well as filled with technobabble and legalese, causing users to unknowingly…

计算与语言 · 计算机科学 2026-05-01 Pengyun Zhu , Qiheng Sun , Long Wen , Yanbo Wang , Yang Cao , Junxu Liu , Deyi Xiong , Jinfei Liu , Zhibo Wang , Kui Ren

Retrieval-Augmented Generation (RAG) enhances the utility of Large Language Models (LLMs) by retrieving external documents. Since the knowledge databases in RAG are predominantly utilized via cloud services, private data in sensitive…

密码学与安全 · 计算机科学 2026-05-29 Xinyuan Zhu , Zekun Fei , Enye Wang , Ruiqi He , Jia Guo , Ruijie Wang , Zheli Liu , Qingkai Zeng

The VoicePrivacy Challenge aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a…

The entity resolution problem requires finding pairs across datasets that belong to different owners but refer to the same entity in the real world. To train and evaluate solutions (either rule-based or machine-learning-based) to the entity…

信息检索 · 计算机科学 2025-06-05 Yixiang Yao , Weizhao Jin , Srivatsan Ravi

Metric Differential Privacy is a generalization of differential privacy tailored to address the unique challenges of text-to-text privatization. By adding noise to the representation of words in the geometric space of embeddings, words are…

计算与语言 · 计算机科学 2023-06-05 Stefan Arnold , Dilara Yesilbas , Sven Weinzierl

Text sanitization, which employs differential privacy to replace sensitive tokens with new ones, represents a significant technique for privacy protection. Typically, its performance in preserving privacy is evaluated by measuring the…

密码学与安全 · 计算机科学 2025-05-06 Meng Tong , Kejiang Chen , Xiaojian Yuan , Jiayang Liu , Weiming Zhang , Nenghai Yu , Jie Zhang

In this document, we present a state of the art of anonymization techniques for classical tabular datasets. This article is geared towards a general public having some knowledge of mathematics and computer science, but with no need for…

密码学与安全 · 计算机科学 2020-01-09 Benjamin Nguyen , Claude Castelluccia

Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation identification, spam detection and content moderation, etc. Despite significant advances in large…

计算与语言 · 计算机科学 2025-07-17 Feng Xiao , Jicong Fan

Named Entity Recognition (NER) has been mostly studied in the context of written text. Specifically, NER is an important step in de-identification (de-ID) of medical records, many of which are recorded conversations between a patient and a…

Existing privacy-preserving speech representation learning methods target a single application domain. In this paper, we present a novel framework to anonymize utterance-level speech embeddings generated by pre-trained encoders and show its…

音频与语音处理 · 电气工程与系统科学 2023-10-27 Minh Tran , Mohammad Soleymani

In recent years, the increasing availability of personal data has raised concerns regarding privacy and security. One of the critical processes to address these concerns is data anonymization, which aims to protect individual privacy and…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Fabio Hellmann , Silvan Mertes , Mohamed Benouis , Alexander Hustinx , Tzung-Chien Hsieh , Cristina Conati , Peter Krawitz , Elisabeth André

Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting…

计算与语言 · 计算机科学 2023-12-22 Xinyi He , Mengyu Zhou , Xinrun Xu , Xiaojun Ma , Rui Ding , Lun Du , Yan Gao , Ran Jia , Xu Chen , Shi Han , Zejian Yuan , Dongmei Zhang

Voice anonymisation can be used to help protect speaker privacy when speech data is shared with untrusted others. In most practical applications, while the voice identity should be sanitised, other attributes such as the spoken content…

音频与语音处理 · 电气工程与系统科学 2024-08-09 Michele Panariello , Massimiliano Todisco , Nicholas Evans

This study addresses the challenge of creating datasets for cybercrime analysis while complying with the requirements of regulations such as the General Data Protection Regulation (GDPR) and Organic Law 10/1995 of the Penal Code. To this…

机器学习 · 计算机科学 2026-04-13 Carlos Jimeno Miguel , Raul Orduna , Francesco Zola

Embeddings, which compress information in raw text into semantics-preserving low-dimensional vectors, have been widely adopted for their efficacy. However, recent research has shown that embeddings can potentially leak private information…

计算与语言 · 计算机科学 2022-10-07 Garam Lee , Minsoo Kim , Jai Hyun Park , Seung-won Hwang , Jung Hee Cheon

Suicidal ideation detection is critical for real-time suicide prevention, yet its progress faces two under-explored challenges: limited language coverage and unreliable annotation practices. Most available datasets are in English, but even…

计算与语言 · 计算机科学 2025-07-22 Amina Dzafic , Merve Kavut , Ulya Bayram

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

Text summarization aims to extract essential information from a piece of text and transform the text into a concise version. Existing unsupervised abstractive summarization models leverage recurrent neural networks framework while the…

计算与语言 · 计算机科学 2020-10-20 Ziyi Yang , Chenguang Zhu , Robert Gmyr , Michael Zeng , Xuedong Huang , Eric Darve