中文
相关论文

相关论文: Textwash -- automated open-source text anonymisati…

200 篇论文

Automatic evaluation of various text quality criteria produced by data-driven intelligent methods is very common and useful because it is cheap, fast, and usually yields repeatable results. In this paper, we present an attempt to automate…

计算与语言 · 计算机科学 2020-06-08 Erion Çano , Ondřej Bojar

For new participants - Executive summary: (1) The task is to develop a voice anonymization system for speech data which conceals the speaker's voice identity while protecting linguistic content, paralinguistic attributes, intelligibility…

In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to…

人机交互 · 计算机科学 2024-12-20 Isadora Krsek , Anubha Kabra , Yao Dou , Tarek Naous , Laura A. Dabbish , Alan Ritter , Wei Xu , Sauvik Das

Organizations are collecting vast amounts of data, but they often lack the capabilities needed to fully extract insights. As a result, they increasingly share data with external experts, such as analysts or researchers, to gain value from…

机器学习 · 计算机科学 2025-05-16 Yusi Wei , Hande Y. Benson , Joseph K. Agor , Muge Capan

Qualitative research often contains personal, contextual, and organizational details that pose privacy risks if not handled appropriately. Manual anonymization is time-consuming, inconsistent, and frequently omits critical identifiers.…

人工智能 · 计算机科学 2026-01-22 Aisvarya Adeseye , Jouni Isoaho , Seppo Virtanen , Mohammad Tahir

This paper proposes a sensor data anonymization model that is trained on decentralized data and strikes a desirable trade-off between data utility and privacy, even in heterogeneous settings where the sensor data have different underlying…

机器学习 · 计算机科学 2023-10-24 Xin Yang , Omid Ardakanian

The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (PHI). Anonymized data…

计算与语言 · 计算机科学 2024-12-17 Murat Gunay , Bunyamin Keles , Raife Hizlan

In this study, we examined the possibility to extract personality traits from a text. We created an extensive dataset by having experts annotate personality traits in a large number of texts from multiple online sources. From these…

计算与语言 · 计算机科学 2019-10-23 Nazar Akrami , Johan Fernquist , Tim Isbister , Lisa Kaati , Björn Pelzer

It is important to study the risks of publishing privacy-sensitive data. Even if sensitive identities (e.g., name, social security number) were removed and advanced data perturbation techniques were applied, several de-anonymization attacks…

社会与信息网络 · 计算机科学 2018-01-18 Wei-Han Lee , Changchang Liu , Shouling Ji , Prateek Mittal , Ruby Lee

In this paper, a new mathematical formulation for the problem of de-anonymizing social network users by actively querying their membership in social network groups is introduced. In this formulation, the attacker has access to a noisy…

信息论 · 计算机科学 2017-10-12 Farhad Shirani , Siddharth Garg , Elza Erkip

In the era of big and ubiquitous data, professionals and students alike are finding themselves needing to perform a number of textual analysis tasks. Historically, the general lack of statistical expertise and programming skills has stopped…

数字图书馆 · 计算机科学 2024-10-30 Faizhal Arif Santosa , Manika Lamba , Crissandra George , J. Stephen Downie

The increasing prevalence of computer vision applications necessitates handling vast amounts of visual data, often containing personal information. While this technology offers significant benefits, it should not compromise privacy. Data…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Mustafa İzzet Muştu , Hazım Kemal Ekenel

We present a method for generating synthetic versions of Twitter data using neural generative models. The goal is protecting individuals in the source data from stylometric re-identification attacks while still releasing data that carries…

计算与语言 · 计算机科学 2018-05-31 Alexander G. Ororbia , Fridolin Linder , Joshua Snoke

Adolescent suicide is a critical global health issue, and speech provides a cost-effective modality for automatic suicide risk detection. Given the vulnerable population, protecting speaker identity is particularly important, as speech…

音频与语音处理 · 电气工程与系统科学 2026-01-26 Ziyun Cui , Sike Jia , Yang Lin , Yinan Duan , Diyang Qu , Runsen Chen , Chao Zhang , Chang Lei , Wen Wu

We argue that governments should mandate a three-tier anonymity framework on social-media platforms as a reactionary measure prompted by the ease-of-production of deepfakes and large-language-model-driven misinformation. The tiers are…

社会与信息网络 · 计算机科学 2026-02-11 David Khachaturov , Roxanne Schnyder , Robert Mullins

With current technology, a number of entities have access to user mobility traces at different levels of spatio-temporal granularity. At the same time, users frequently reveal their location through different means, including geo-tagged…

密码学与安全 · 计算机科学 2018-11-16 Apostolos Pyrgelis , Nicolas Kourtellis , Ilias Leontiadis , Joan Serrà , Claudio Soriente

Attribute inference - the process of analyzing publicly available data in order to uncover hidden information - has become a major threat to privacy, given the recent technological leap in machine learning. One way to tackle this threat is…

人工智能 · 计算机科学 2023-04-25 Marcin Waniek , Navya Suri , Abdullah Zameek , Bedoor AlShebli , Talal Rahwan

The growing use of large language models has increased interest in sharing textual data in a privacy-preserving manner. One prominent line of work addresses this challenge through text rewriting under Local Differential Privacy (LDP), where…

密码学与安全 · 计算机科学 2026-03-25 Weijun Li , Arnaud Grivet Sébert , Qiongkai Xu , Annabelle McIver , Mark Dras

The recent surge in high-quality open-source Generative AI text models (colloquially: LLMs), as well as efficient finetuning techniques, have opened the possibility of creating high-quality personalized models that generate text attuned to…

计算与语言 · 计算机科学 2025-06-03 Eugenia Iofinova , Andrej Jovanovic , Dan Alistarh

Online questionnaires that use crowd-sourcing platforms to recruit participants have become commonplace, due to their ease of use and low costs. Artificial Intelligence (AI) based Large Language Models (LLM) have made it easy for bad actors…

人机交互 · 计算机科学 2024-02-02 Benjamin Lebrun , Sharon Temtsin , Andrew Vonasch , Christoph Bartneck
‹ 上一页 1 8 9 10 下一页 ›