基于众包标注的手写文本识别
计算机视觉与模式识别
2023-06-21 v1 人工智能
摘要
在本文中,我们探索了当存在多个不完美或含噪转写时,训练手写文本识别模型的不同方法。我们考虑了多种训练配置,例如选择单一转写、保留所有转写,或从所有可用标注中计算聚合转写。此外,我们评估了基于质量的数据选择的影响,即从训练集中移除标注者之间一致性较低的样本。我们的实验在法国 Belfort 市 1790 年至 1946 年间的市政登记册上进行。结果表明,计算共识转写或在多个转写上训练是良好的替代方案。然而,基于标注者之间一致程度选择训练样本会在训练数据中引入偏差,并不能改善结果。我们的数据集公开于 Zenodo:https://zenodo.org/record/8041668。
引用
@article{arxiv.2306.10878,
title = {Handwritten Text Recognition from Crowdsourced Annotations},
author = {Solène Tarride and Tristan Faine and Mélodie Boillet and Harold Mouchère and Christopher Kermorvant},
journal= {arXiv preprint arXiv:2306.10878},
year = {2023}
}
备注
Accepted to the 7th International Workshop on Historical Document Imaging and Processing (HIP 23)