中文

挪威国家图书馆萨米文本光学字符识别方法的比较分析

计算与语言 2025-01-14 v1 计算机视觉与模式识别

摘要

光学字符识别(OCR)对于挪威国家图书馆(NLN)的数字化过程至关重要,因为它将扫描文档转换为机器可读文本。然而,对于NLN馆藏中的萨米文档,OCR精度不足。鉴于OCR质量影响下游处理,评估和改进萨米语文本的OCR对于使这些资源可访问是必要的。为满足这一需求,本工作微调并评估了三种成熟的OCR方法——Transkribus、Tesseract和TrOCR——用于转录NLN馆藏中的萨米文本。我们的结果表明,Transkribus和TrOCR在此任务上优于Tesseract,而Tesseract在域外数据集上取得了更优性能。此外,我们表明,微调预训练模型并用机器标注和合成文本图像补充人工标注,即使只有中等数量的人工标注数据,也能为萨米语产生准确的OCR。

关键词

引用

@article{arxiv.2501.07300,
  title  = {Comparative analysis of optical character recognition methods for S\'ami texts from the National Library of Norway},
  author = {Tita Enstad and Trond Trosterud and Marie Iversdatter Røsok and Yngvil Beyer and Marie Roald},
  journal= {arXiv preprint arXiv:2501.07300},
  year   = {2025}
}

备注

To be published in Proceedings of the 25th Nordic Conference on Computational Linguistics (NoDaLiDa)