面向中世纪拉丁文的手写文本识别定制系统
计算机视觉与模式识别
2023-08-21 v1 计算与语言
计算机与社会
机器学习
机器学习
摘要
巴伐利亚科学与人文学院旨在将其《中世纪拉丁文词典》数字化。该词典包含指向中世纪拉丁文(一种低资源语言)词目的记录卡。数字化过程的关键一步是对这些记录卡上手写词目进行手写文本识别(HTR)。在我们的工作中,我们引入了一个端到端流水线,专为中世纪拉丁文词典定制,用于定位、提取和转写词目。我们采用两个最先进(SOTA)的图像分割模型为 HTR 任务准备初始数据集。此外,我们尝试不同的基于 transformer 的模型,并进行一组实验以探索不同视觉编码器与 GPT-2 解码器组合的能力。另外,我们还应用了广泛的数据增强,从而得到一个极具竞争力的模型。性能最佳的配置实现了 0.015 的字符错误率(CER),甚至优于商业的 Google Cloud Vision 模型,并表现出更稳定的性能。
引用
@article{arxiv.2308.09368,
title = {A tailored Handwritten-Text-Recognition System for Medieval Latin},
author = {Philipp Koch and Gilary Vera Nuñez and Esteban Garces Arias and Christian Heumann and Matthias Schöffel and Alexander Häberlin and Matthias Aßenmacher},
journal= {arXiv preprint arXiv:2308.09368},
year = {2023}
}
备注
This paper has been accepted at the First Workshop on Ancient Language Processing, co-located with RANLP 2023. This is the author's version of the work. The definite version of record will be published in the proceedings