English

From OCR to Analysis: Tracking Correction Provenance in Digital Humanities Pipelines

Human-Computer Interaction 2026-05-07 v4

Abstract

Optical Character Recognition (OCR) is a critical but error-prone stage in digital humanities text pipelines. While OCR correction improves usability for downstream NLP tasks, common workflows often overwrite intermediate decisions, obscuring how textual transformations affect scholarly interpretation. We present a provenance-aware framework for OCR-corrected humanities corpora that records correction lineage at the span level, including edit type, correction source, confidence, and revision status. Using a pilot corpus of historical texts, we compare downstream named entity extraction across raw OCR, fully corrected text, and provenance-filtered corrections. Our results show that correction pathways can substantially alter extracted entities and document-level interpretations, while provenance signals help identify unstable outputs and prioritize human review. We argue that provenance should be treated as a first-class analytical layer in NLP for digital humanities, supporting reproducibility, source criticism, and uncertainty-aware interpretation.

Keywords

Cite

@article{arxiv.2603.00884,
  title  = {From OCR to Analysis: Tracking Correction Provenance in Digital Humanities Pipelines},
  author = {Haoze Guo and Ziqi Wei},
  journal= {arXiv preprint arXiv:2603.00884},
  year   = {2026}
}

Comments

In Proceedings of the 6th International Conference on Natural Language Processing for Digital Humanities

R2 v1 2026-07-01T10:57:37.983Z