消歧数字序列以破译古代会计语料库
计算与语言
2025-04-29 v2
摘要
记数系统将抽象的数字量编码为具体的书写字符序列。现代文字所使用的记数系统往往精确且无歧义,但对于古代且部分破译的原始埃兰(PE)文字而言并非如此,其书面数字根据读取时所使用的系统,最多可有四种不同的读法。我们考虑的任务是消歧这些读法,以确定该语料库中记录的数字量的值。我们通过算法为每个 PE 数字符号提取了一份可能的读法列表,并贡献了两种基于原始文档结构属性和通过自举算法学习的分类器的消歧技术。我们还贡献了一个用于评估消歧技术的测试集,以及一种用于自举分类器的谨慎规则选择新方法。我们的分析证实了关于这种文字的现有直觉,并揭示了泥板内容与数字大小之间先前未知的相关性。这项工作对于理解和破译 PE 至关重要,因为该语料库以会计记录为主,包含的数字标记远多于文本标记。
引用
@article{arxiv.2502.00090,
title = {Disambiguating Numeral Sequences to Decipher Ancient Accounting Corpora},
author = {Logan Born and M. Willis Monroe and Kathryn Kelley and Anoop Sarkar},
journal= {arXiv preprint arXiv:2502.00090},
year = {2025}
}
备注
Englund 1996 incorrectly reported the relative values of signs in the decimal system. An earlier version of this paper used those values. This update fixes those mistakes and retrains our models using the corrected readings. Our analysis and discussion remain similar to the original, but the performance of the baseline model is now stronger