English

Entities, Dates, and Languages: Zero-Shot on Historical Texts with T0

Computation and Language 2022-04-12 v1

Abstract

In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods. Using a historical newspaper corpus in 3 languages as test-bed, we use prompts to extract possible named entities. Our results show that a naive approach for prompt-based zero-shot multilingual Named Entity Recognition is error-prone, but highlights the potential of such an approach for historical languages lacking labeled datasets. Moreover, we also find that T0-like models can be probed to predict the publication date and language of a document, which could be very relevant for the study of historical texts.

Cite

@article{arxiv.2204.05211,
  title  = {Entities, Dates, and Languages: Zero-Shot on Historical Texts with T0},
  author = {Francesco De Toni and Christopher Akiki and Javier de la Rosa and Clémentine Fourrier and Enrique Manjavacas and Stefan Schweter and Daniel van Strien},
  journal= {arXiv preprint arXiv:2204.05211},
  year   = {2022}
}
R2 v1 2026-06-24T10:44:42.193Z