English

De-identification of Unstructured Clinical Texts from Sequence to Sequence Perspective

Computation and Language 2021-09-13 v2 Cryptography and Security Machine Learning

Abstract

In this work, we propose a novel problem formulation for de-identification of unstructured clinical text. We formulate the de-identification problem as a sequence to sequence learning problem instead of a token classification problem. Our approach is inspired by the recent state-of -the-art performance of sequence to sequence learning models for named entity recognition. Early experimentation of our proposed approach achieved 98.91% recall rate on i2b2 dataset. This performance is comparable to current state-of-the-art models for unstructured clinical text de-identification.

Keywords

Cite

@article{arxiv.2108.07971,
  title  = {De-identification of Unstructured Clinical Texts from Sequence to Sequence Perspective},
  author = {Md Monowar Anjum and Noman Mohammed and Xiaoqian Jiang},
  journal= {arXiv preprint arXiv:2108.07971},
  year   = {2021}
}

Comments

Accepted in Poster Track for ACM CCS 2021

R2 v1 2026-06-24T05:12:40.102Z