English

From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction

Computation and Language 2020-07-23 v1

Abstract

A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a great part of the automatic approaches have been relying on rules or supervised machine learning. We present a fully automatic unsupervised way of extracting parallel data for training a character-based sequence-to-sequence NMT (neural machine translation) model to conduct OCR error correction.

Keywords

Cite

@article{arxiv.1910.05535,
  title  = {From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction},
  author = {Mika Hämäläinen and Simon Hengchen},
  journal= {arXiv preprint arXiv:1910.05535},
  year   = {2020}
}