中文

阅读不可读之物:使用图像到文本语言模型创建 19 世纪英语报纸数据集

计算与语言 2025-02-24 v1 数字图书馆 机器学习

摘要

奥斯卡·怀特晚晚曾说:“文学与新闻的区别在于,新闻不可读,而文学却不被阅读。”不幸的是,奥斯卡·怀特晚晚 19 世纪的数字存档新闻常常没有或遗传质量差的光学字符识别(OCR),降低了这些档案的可访问性,使其在象征性和字面意义上都变得不可读。本文通过对“十九世纪连载作品选集”(NCSE)进行 OCR 处理来解决此问题,该集合包含 84,000 页的 19 世纪英语报纸和期刊,采用 Pixtral 12B,这是一种预训练的图像到文本语言模型。将 Pixtral 的 OCR 能力与 4 种其他 OCR 方法进行比较, achieves a median character error rate of 1%,是下一名模型的 5 倍。 resulting NCSE v2.0 dataset features improved article identification, high-quality OCR, and text classified into four types and seventeen topics. The dataset contains 1.4 million entries, and 321 million words. Example use cases demonstrate analysis of topic similarity, readability, and event tracking. NCSE v2.0 is freely available to encourage historical and sociological research. As a result, 21st-century readers can now share Oscar Wilde's disappointment with 19th-century journalistic standards, reading the unreadable from the comfort of their own computers.

关键词

引用

@article{arxiv.2502.14901,
  title  = {Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models},
  author = {Jonathan Bourne},
  journal= {arXiv preprint arXiv:2502.14901},
  year   = {2025}
}