PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues
Abstract
Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds and the recall/precision trade-off still limit the uptake of these anonymization pipelines. We present PIIvot, a lighter-weight framework for PII anonymization that leverages knowledge of the data context to simplify the PII detection problem. To demonstrate its effectiveness, we also contribute QATD-2k, the largest open-source real-world tutoring dataset of its kind, to support the demand for quality educational dialogue data.
Keywords
Cite
@article{arxiv.2505.16931,
title = {PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues},
author = {Matthew Zent and Digory Smith and Simon Woodhead},
journal= {arXiv preprint arXiv:2505.16931},
year = {2025}
}
Comments
6 pages, 2 figures, submitted to EMNLP 2025, for associated dataset, see https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k