中文

过去的西班牙文本:Transkribus、Tesseract 与 Granite 的实验

计算机视觉与模式识别 2025-07-08 v1 计算与语言

摘要

本文介绍了 GRESEL 团队在 IberLEF 2025 共享任务 PastReader: Transcribing Texts from the Past 中的实验与结果。 conducted 了三类实验,旨在参与该任务并实现不同方法之间的比较。这些包括使用基于网页的 OCR 服务、传统 OCR 引擎和紧凑型多模态模型。所有实验都在消费级硬件上运行,尽管缺乏高性能计算能力,但提供了足够的存储和稳定性。结果虽令人满意,但仍有改进空间。未来的工作将聚焦于利用该共享任务提供的西班牙语数据集,结合西班牙国家图书馆(BNE)合作,探索新的技术和想法。

关键词

引用

@article{arxiv.2507.04878,
  title  = {Transcribing Spanish Texts from the Past: Experiments with Transkribus, Tesseract and Granite},
  author = {Yanco Amor Torterolo-Orta and Jaione Macicior-Mitxelena and Marina Miguez-Lamanuzzi and Ana García-Serrano},
  journal= {arXiv preprint arXiv:2507.04878},
  year   = {2025}
}

备注

This paper was written as part of a shared task organized within the 2025 edition of the Iberian Languages Evaluation Forum (IberLEF 2025), held at SEPLN 2025 in Zaragoza. This paper describes the joint participation of two teams in said competition, GRESEL1 and GRESEL2, each with an individual paper that will be published in CEUR