English
Related papers

Related papers: Anonpsy: A Graph-Based Framework for Structure-Pre…

200 papers

Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions like sepsis. While unstructured clinical narratives offer semantically rich and contextually…

Computation and Language · Computer Science 2026-05-15 Sayantan Kumar , Shahriar Noroozizadeh , Juyong Kim , Jeremy C. Weiss

Given a pool of observations selected from a sensor stream, input data can be robustly represented, via a multiscale process, in terms of invariant concepts, and themes. Applying this to episodic natural language data, one may obtain a…

Artificial Intelligence · Computer Science 2020-10-19 Mark Burgess

The increasing capabilities of deep neural networks for re-identification, combined with the rise in public surveillance in recent years, pose a substantial threat to individual privacy. Event cameras were initially considered as a…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Katharina Bendig , René Schuster , Nicole Thiemer , Karen Joisten , Didier Stricker

This work presents AnonLFI 2.0, a modular pseudonymization framework for CSIRTs that uses HMAC SHA256 to generate strong and reversible pseudonyms, preserves XML and JSON structures, and integrates OCR and technical recognizers for PII and…

Cryptography and Security · Computer Science 2025-11-21 Cristhian Kapelinski , Douglas Lautert , Beatriz Machado , Diego Kreutz

This work investigates the effectiveness of different pseudonymization techniques, ranging from rule-based substitutions to using pre-trained Large Language Models (LLMs), on a variety of datasets and models used for two widely used NLP…

Computation and Language · Computer Science 2023-06-12 Oleksandr Yermilov , Vipul Raheja , Artem Chernodub

The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in…

Computation and Language · Computer Science 2023-03-24 Xavier Tannier , Perceval Wajsbürt , Alice Calliger , Basile Dura , Alexandre Mouchet , Martin Hilka , Romain Bey

Privacy protection of medical image data is challenging. Even if metadata is removed, brain scans are vulnerable to attacks that match renderings of the face to facial image databases. Solutions have been developed to de-identify diagnostic…

Image and Video Processing · Electrical Eng. & Systems 2021-10-20 Lennart Alexander Van der Goten , Tobias Hepp , Zeynep Akata , Kevin Smith

Clinical text processing has gained more and more attention in recent years. The access to sensitive patient data, on the other hand, is still a big challenge, as text cannot be shared without legal hurdles and without removing personal…

Computation and Language · Computer Science 2022-09-02 Iyadh Ben Cheikh Larbi , Aljoscha Burchardt , Roland Roller

Besides its linguistic content, our speech is rich in biometric information that can be inferred by classifiers. Learning privacy-preserving representations for speech signals enables downstream tasks without sharing unnecessary, private…

Sound · Computer Science 2021-06-18 Dimitrios Stoidis , Andrea Cavallaro

Graphs are ubiquitous structures found in numerous real-world applications, such as drug discovery, recommender systems, and social network analysis. To model graph-structured data, graph neural networks (GNNs) have become a popular tool.…

Computation and Language · Computer Science 2025-02-10 Zehong Wang , Sidney Liu , Zheyuan Zhang , Tianyi Ma , Chuxu Zhang , Yanfang Ye

The increasing popularity of social networks has initiated a fertile research area in information extraction and data mining. Anonymization of these social graphs is important to facilitate publishing these data sets for analysis by…

Databases · Computer Science 2015-03-13 Sudipto Das , Omer Egecioglu , Amr El Abbadi

Unstructured text from legal, medical, and administrative sources offers a rich but underutilized resource for research in public health and the social sciences. However, large-scale analysis is hampered by two key challenges: the presence…

Computation and Language · Computer Science 2025-07-16 Anders Ledberg , Anna Thalén

Automated clinical text anonymization has the potential to unlock the widespread sharing of textual health data for secondary usage while assuring patient privacy and safety. Despite the proposal of many complex and theoretically successful…

Computation and Language · Computer Science 2024-06-05 David Pissarra , Isabel Curioso , João Alveira , Duarte Pereira , Bruno Ribeiro , Tomás Souper , Vasco Gomes , André V. Carreiro , Vitor Rolla

Advances in large language models (LLMs) have enabled a wide range of applications. However, depression prediction is hindered by the lack of large-scale, high-quality, and rigorously annotated datasets. This study introduces DepressLLM,…

Computation and Language · Computer Science 2025-08-13 Sehwan Moon , Aram Lee , Jeong Eun Kim , Hee-Ju Kang , Il-Seon Shin , Sung-Wan Kim , Jae-Min Kim , Min Jhon , Ju-Wan Kim

In many countries, personal information that can be published or shared between organizations is regulated and, therefore, documents must undergo a process of de-identification to eliminate or obfuscate confidential data. Our work focuses…

Computation and Language · Computer Science 2019-10-10 Diego Garat , Dina Wonsever

Due to the rapid advancement of Large Language Model (LLM), the whole community eagerly consumes any available text data in order to train the LLM. Currently, large portion of the available text data are collected from internet, which has…

Artificial Intelligence · Computer Science 2024-06-21 Ya-Lun Li

Hallucinations, the generation of apparently convincing yet false statements, remain a major barrier to the safe deployment of LLMs. Building on the strong performance of self-detection methods, we examine the use of structured knowledge…

Computation and Language · Computer Science 2025-12-30 Sahil Kale , Antonio Luca Alfeo

Protecting patient privacy in healthcare records is a top priority, and redaction is a commonly used method for obscuring directly identifiable information in text. Rule-based methods have been widely used, but their precision is often low…

Sharing sensitive texts for scientific purposes requires appropriate techniques to protect the privacy of patients and healthcare personnel. Anonymizing textual data is particularly challenging due to the presence of diverse unstructured…

Computation and Language · Computer Science 2025-02-20 Ibrahim Baroud , Lisa Raithel , Sebastian Möller , Roland Roller

Medical imaging has significantly advanced computer-aided diagnosis, yet its re-identification (ReID) risks raise critical privacy concerns, calling for de-identification (DeID) techniques. Unfortunately, existing DeID methods neither…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yuan Tian , Shuo Wang , Rongzhao Zhang , Zijian Chen , Yankai Jiang , Chunyi Li , Xiangyang Zhu , Fang Yan , Qiang Hu , XiaoSong Wang , Guangtao Zhai