基于arXiv的大规模合成数据用于历史科学文献的OCR后校正
数字图书馆
2023-09-22 v1 天体物理仪器与方法
摘要
在“数字化时代”(约1997年)之前发表的科学文章需要光学字符识别(OCR)将扫描文档转换为机器可读文本,这一过程常产生错误。我们开发了一条生成合成真值/OCR数据集的流程,以校正NASA天体物理数据系统(ADS)天体物理文献馆藏的OCR结果。通过挖掘arXiv,我们创建了——据作者所知——最大的科学合成真值/OCR后校正数据集,包含203,354,393个字符对。我们提供了用该数据集训练的基线模型,并发现历史OCR文本的字符错误率和词错误率平均分别改善了7.71%和18.82%。当用于分类句子中的行内数学部分时,我们获得77.82%的分类F1分数。用于探索该数据集的交互式仪表盘可在线获取:https://readingtimemachine.github.io/projects/1-ocr-groundtruth-may2023,且在arXiv协议限制内数据与代码托管于GitHub:https://github.com/ReadingTimeMachine/ocr_post_correction。
引用
@article{arxiv.2309.11549,
title = {Large Synthetic Data from the arXiv for OCR Post Correction of Historic Scientific Articles},
author = {Jill P. Naiman and Morgan G. Cosillo and Peter K. G. Williams and Alyssa Goodman},
journal= {arXiv preprint arXiv:2309.11549},
year = {2023}
}
备注
6 pages, 1 figure, 1 table; training/validation/test datasets and all model weights to be linked on Zenodo on publication