English
Related papers

Related papers: EDDA-Coordinata: An Annotated Dataset of Historica…

200 papers

Diderot's \textit{Encyclop\'edie} is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era. \textit{Wikipedia} has the same ambition with a much greater scope. However, the lack of digital…

Computation and Language · Computer Science 2024-06-06 Pierre Nugues

Subnational location data of disaster events are critical for risk assessment and disaster risk reduction. Disaster databases such as EM-DAT often report locations in unstructured textual form, with inconsistent granularity or spelling,…

Artificial Intelligence · Computer Science 2025-11-20 Michele Ronco , Damien Delforge , Wiebke S. Jäger , Christina Corbane

We study the problem of resolving a perhaps misspelled address of a location into geographic coordinates of latitude and longitude. Our data structure solves this problem within a few milliseconds even for misspelled and fragmentary…

Information Retrieval · Computer Science 2011-02-17 Christian Jung , Daniel Karch , Sebastian Knopp , Dennis Luxen , Peter Sanders

Automatically extracting the geometric content from the hundreds of thousands of diagrams drawn in historical manuscripts would enable historians to study the diffusion of astronomical knowledge on a global scale. However, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Syrine Kalleli , Scott Trigg , Ségolène Albouy , Mathieu Husson , Mathieu Aubry

Historical maps offer an invaluable perspective into territory evolution across past centuries--long before satellite or remote sensing technologies existed. Deep learning methods have shown promising results in segmenting historical maps,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Marta López-Rauhut , Hongyu Zhou , Mathieu Aubry , Loic Landrieu

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

Information Retrieval · Computer Science 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu

Currently available grammatical error correction (GEC) datasets are compiled using well-formed written text, limiting the applicability of these datasets to other domains such as informal writing and dialog. In this paper, we present a…

Computation and Language · Computer Science 2025-08-27 Xun Yuan , Derek Pham , Sam Davidson , Zhou Yu

Named Entity Recognition (NER) in historical texts presents unique challenges due to non-standardized language, archaic orthography, and nested or overlapping entities. This study benchmarks a diverse set of NER approaches, ranging from…

Computation and Language · Computer Science 2025-06-04 Ludovic Moncla , Hédi Zeghidi

Archivists, textual scholars, and historians often produce digital editions of historical documents. Using markup schemes such as those of the Text Encoding Initiative and EpiDoc, these digital editions often record documents' semantic…

Computer Vision and Pattern Recognition · Computer Science 2021-12-24 Alejandro H. Toselli , Si Wu , David A. Smith

Extracting the "correct" location information from text data, i.e., determining the place of event, has long been a goal for automated text processing. To approximate human-like coding schema, we introduce a supervised machine learning…

Computation and Language · Computer Science 2019-08-28 Sophie J. Lee , Howard Liu , Michael D. Ward

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific…

Digital Libraries · Computer Science 2026-01-09 Zhiyin Tan , Changxu Duan

We present EDA: easy data augmentation techniques for boosting performance on text classification tasks. EDA consists of four simple but powerful operations: synonym replacement, random insertion, random swap, and random deletion. On five…

Computation and Language · Computer Science 2019-08-27 Jason Wei , Kai Zou

Diagram parsing is an important foundation for geometry problem solving, attracting increasing attention in the field of intelligent education and document image understanding. Due to the complex layout and between-primitive relationship,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-23 Yihan Hao , Mingliang Zhang , Fei Yin , Linlin Huang

This paper proposes AEDA (An Easier Data Augmentation) technique to help improve the performance on text classification tasks. AEDA includes only random insertion of punctuation marks into the original text. This is an easier technique to…

Computation and Language · Computer Science 2021-08-31 Akbar Karimi , Leonardo Rossi , Andrea Prati

The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisation. Image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Shuo Wang , Yan Zhu , Xiaoyuan Luo , Zhiwei Yang , Yizhe Zhang , Peiyao Fu , Manning Wang , Zhijian Song , Quanlin Li , Pinghong Zhou , Yike Guo

Matching place names across writing systems is a persistent obstacle to the integration of multilingual geographic sources, whether modern gazetteers, medieval itineraries, or colonial-era surveys. Existing approaches depend on…

Computation and Language · Computer Science 2026-03-31 Stephen Gadd

It is becoming common to archive research datasets that are not only large but also numerous. In addition, their corresponding metadata and the software required to analyse or display them need to be archived. Yet the manual curation of…

Digital Libraries · Computer Science 2011-08-24 Daniel Lemire , Andre Vellino

The latest developments in digital have provided large data sets that can increasingly easily be accessed and used. These data sets often contain indirect localisation information, such as historical addresses. Historical geocoding is the…

Databases · Computer Science 2018-06-01 Rémi Cura , Bertrand Dumenieu , Nathalie Abadie , Benoit Costes , Julien Perret , Maurizio Gribaudi

We propose a global entity disambiguation (ED) model based on BERT. To capture global contextual information for ED, our model treats not only words but also entities as input tokens, and solves the task by sequentially resolving mentions…

Computation and Language · Computer Science 2022-05-03 Ikuya Yamada , Koki Washio , Hiroyuki Shindo , Yuji Matsumoto

Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we developed a version…

Computation and Language · Computer Science 2026-03-12 Masato Kikuchi , Masatsugu Ono , Toshioki Soga , Tetsu Tanabe , Tadachika Ozono
‹ Prev 1 2 3 10 Next ›