English

Cross-Domain Transfer and Few-Shot Learning for Personal Identifiable Information Recognition

Computation and Language 2026-01-13 v3

Abstract

Accurate recognition of personally identifiable information (PII) is central to automated text anonymization. This paper investigates the effectiveness of cross-domain model transfer, multi-domain data fusion, and sample-efficient learning for PII recognition. Using annotated corpora from healthcare (I2B2), legal (TAB), and biography (Wikipedia), we evaluate models across four dimensions: in-domain performance, cross-domain transferability, fusion, and few-shot learning. Results show legal-domain data transfers well to biographical texts, while medical domains resist incoming transfer. Fusion benefits are domain-specific, and high-quality recognition is achievable with only 10% of training data in low-specialization domains.

Keywords

Cite

@article{arxiv.2507.11862,
  title  = {Cross-Domain Transfer and Few-Shot Learning for Personal Identifiable Information Recognition},
  author = {Junhong Ye and Xu Yuan and Xinying Qiu},
  journal= {arXiv preprint arXiv:2507.11862},
  year   = {2026}
}

Comments

Accepted to CLNLP 2025