中文
相关论文

相关论文: Disambiguating Numeral Sequences to Decipher Ancie…

200 篇论文

Byte Pair Encoding (BPE) tokenizers, widely used in Large Language Models, face challenges in multilingual settings, including penalization of non-Western scripts and the creation of tokens with partial UTF-8 sequences. Pretokenization,…

计算与语言 · 计算机科学 2025-06-02 Sander Land , Catherine Arnett

The complex Ancient Egyptian (AE) writing system was characterised by widespread use of graphemic classifiers (determinatives): silent (unpronounced) hieroglyphic signs clarifying the meaning or indicating the pronunciation of the host…

计算与语言 · 计算机科学 2024-07-02 Dmitry Nikolaev , Jorke Grotenhuis , Haleli Harel , Orly Goldwasser

In this paper we examine the benefit of performing named entity recognition (NER) and co-reference resolution to an English and a Greek corpus used for text segmentation. The aim here is to examine whether the combination of text…

计算与语言 · 计算机科学 2016-10-31 Pavlina Fragkou

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus of 3.6M sentences…

计算与语言 · 计算机科学 2023-06-08 Vít Novotný , Kristýna Luger , Michal Štefánik , Tereza Vrabcová , Aleš Horák

Numeral systems across the world's languages vary in fascinating ways, both regarding their synchronic structure and the diachronic processes that determined how they evolved in their current shape. For a proper comparison of numeral…

Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks…

计算与语言 · 计算机科学 2026-05-29 Pierre-Antoine Lequeu , Camille Barboule , Benjamin Piwowarski

Ancient scripts, e.g., Egyptian hieroglyphs, Oracle Bone Inscriptions, and Ancient Greek inscriptions, serve as vital carriers of human civilization, embedding invaluable historical and cultural information. Automating ancient script image…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Xiaolei Diao , Rite Bo , Yanling Xiao , Lida Shi , Zhihan Zhou , Hao Xu , Chuntao Li , Xiongfeng Tang , Massimo Poesio , Cédric M. John , Daqian Shi

This paper describes some of the recent work of project AMALGAM (automatic mapping among lexico-grammatical annotation models). We are investigating ways to map between the leading corpus annotation schemes in order to improve their…

cmp-lg · 计算机科学 2008-02-03 John Hughes , Clive Souter , Eric Atwell

Pebbles (calculos in Latin) are the "bits" used in the Ancients' four function calculator / computer. The Ancient Computer's normal mode is to work with numbers in what we would call exponential notation. Decimal numbers can have up to 10…

历史与综述 · 数学 2012-06-21 Stephen Kent Stephenson

Name Entity Disambiguation is the Natural Language Processing task of identifying textual records corresponding to the same Named Entity, i.e. real-world entities represented as a list of attributes (names, places, organisations, etc.). In…

Distant labeling for information extraction (IE) suffers from noisy training data. We describe a way of reducing the noise associated with distant IE by identifying coupling constraints between potential instance labels. As one example of…

计算与语言 · 计算机科学 2016-01-05 Lidong Bing , Mingyang Ling , Richard C. Wang , William W. Cohen

We consider the verification of parameterized networks of replicated processes whose architecture is described by hyperedge-replacement graph grammars. Due to the undecidability of verification problems such as reachability or coverability…

形式语言与自动机理论 · 计算机科学 2025-02-24 Marius Bozga , Radu Iosif , Arnaud Sangnier , Neven Villani

EHR systems lack a unified code system forrepresenting medical concepts, which acts asa barrier for the deployment of deep learningmodels in large scale to multiple clinics and hos-pitals. To overcome this problem, we…

计算与语言 · 计算机科学 2022-01-19 Kyunghoon Hur , Jiyoung Lee , Jungwoo Oh , Wesley Price , Young-Hak Kim , Edward Choi

Deciphering historical substitution ciphers is a challenging problem. Example problems that have been previously studied include detecting cipher type, detecting plaintext language, and acquiring the substitution key for segmented ciphers.…

计算与语言 · 计算机科学 2022-05-26 Nada Aldarrab , Jonathan May

Substantial increase in the use of Electronic Health Records (EHRs) has opened new frontiers for predictive healthcare. However, while EHR systems are nearly ubiquitous, they lack a unified code system for representing medical concepts.…

机器学习 · 计算机科学 2022-03-21 Kyunghoon Hur , Jiyoung Lee , Jungwoo Oh , Wesley Price , Young-Hak Kim , Edward Choi

This work introduces an anonymization scheme for a corpus of texts to safeguard metadata from disclosure. It specifically aims to prevent large language models from identifying metadata associated with texts, thereby avoiding their…

应用统计 · 统计学 2025-05-28 Jan Greve , Lukas Sablica

Timely analysis of cyber-security information necessitates automated information extraction from unstructured text. While state-of-the-art extraction methods produce extremely accurate results, they require ample training data, which is…

信息检索 · 计算机科学 2014-06-11 Robert A. Bridges , Corinne L. Jones , Michael D. Iannacone , Kelly M. Testa , John R. Goodall

Abbreviations are unavoidable yet critical parts of the medical text. Using abbreviations, especially in clinical patient notes, can save time and space, protect sensitive information, and help avoid repetitions. However, most abbreviations…

计算与语言 · 计算机科学 2022-10-07 Mucahit Cevik , Sanaz Mohammad Jafari , Mitchell Myers , Savas Yildirim

Tokenisation - "the process of splitting text into atomic parts" (Brezina & Timperley, 2017: 1) - is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring…

计算与语言 · 计算机科学 2025-07-03 Matteo Di Cristofaro

This paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge,…

计算机视觉与模式识别 · 计算机科学 2019-07-02 Michele Alberti , Lars Vögtlin , Vinaychandran Pondenkandath , Mathias Seuret , Rolf Ingold , Marcus Liwicki
‹ 上一页 1 2 3 10 下一页 ›