中文
相关论文

相关论文: EDDA-Coordinata: An Annotated Dataset of Historica…

200 篇论文

Diderot's \textit{Encyclop\'edie} is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era. \textit{Wikipedia} has the same ambition with a much greater scope. However, the lack of digital…

计算与语言 · 计算机科学 2024-06-06 Pierre Nugues

Subnational location data of disaster events are critical for risk assessment and disaster risk reduction. Disaster databases such as EM-DAT often report locations in unstructured textual form, with inconsistent granularity or spelling,…

人工智能 · 计算机科学 2025-11-20 Michele Ronco , Damien Delforge , Wiebke S. Jäger , Christina Corbane

We study the problem of resolving a perhaps misspelled address of a location into geographic coordinates of latitude and longitude. Our data structure solves this problem within a few milliseconds even for misspelled and fragmentary…

信息检索 · 计算机科学 2011-02-17 Christian Jung , Daniel Karch , Sebastian Knopp , Dennis Luxen , Peter Sanders

Automatically extracting the geometric content from the hundreds of thousands of diagrams drawn in historical manuscripts would enable historians to study the diffusion of astronomical knowledge on a global scale. However, state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Syrine Kalleli , Scott Trigg , Ségolène Albouy , Mathieu Husson , Mathieu Aubry

Historical maps offer an invaluable perspective into territory evolution across past centuries--long before satellite or remote sensing technologies existed. Deep learning methods have shown promising results in segmenting historical maps,…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Marta López-Rauhut , Hongyu Zhou , Mathieu Aubry , Loic Landrieu

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

信息检索 · 计算机科学 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu

Currently available grammatical error correction (GEC) datasets are compiled using well-formed written text, limiting the applicability of these datasets to other domains such as informal writing and dialog. In this paper, we present a…

计算与语言 · 计算机科学 2025-08-27 Xun Yuan , Derek Pham , Sam Davidson , Zhou Yu

Named Entity Recognition (NER) in historical texts presents unique challenges due to non-standardized language, archaic orthography, and nested or overlapping entities. This study benchmarks a diverse set of NER approaches, ranging from…

计算与语言 · 计算机科学 2025-06-04 Ludovic Moncla , Hédi Zeghidi

Archivists, textual scholars, and historians often produce digital editions of historical documents. Using markup schemes such as those of the Text Encoding Initiative and EpiDoc, these digital editions often record documents' semantic…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Alejandro H. Toselli , Si Wu , David A. Smith

Extracting the "correct" location information from text data, i.e., determining the place of event, has long been a goal for automated text processing. To approximate human-like coding schema, we introduce a supervised machine learning…

计算与语言 · 计算机科学 2019-08-28 Sophie J. Lee , Howard Liu , Michael D. Ward

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific…

数字图书馆 · 计算机科学 2026-01-09 Zhiyin Tan , Changxu Duan

We present EDA: easy data augmentation techniques for boosting performance on text classification tasks. EDA consists of four simple but powerful operations: synonym replacement, random insertion, random swap, and random deletion. On five…

计算与语言 · 计算机科学 2019-08-27 Jason Wei , Kai Zou

Diagram parsing is an important foundation for geometry problem solving, attracting increasing attention in the field of intelligent education and document image understanding. Due to the complex layout and between-primitive relationship,…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Yihan Hao , Mingliang Zhang , Fei Yin , Linlin Huang

This paper proposes AEDA (An Easier Data Augmentation) technique to help improve the performance on text classification tasks. AEDA includes only random insertion of punctuation marks into the original text. This is an easier technique to…

计算与语言 · 计算机科学 2021-08-31 Akbar Karimi , Leonardo Rossi , Andrea Prati

The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisation. Image-text…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Shuo Wang , Yan Zhu , Xiaoyuan Luo , Zhiwei Yang , Yizhe Zhang , Peiyao Fu , Manning Wang , Zhijian Song , Quanlin Li , Pinghong Zhou , Yike Guo

Matching place names across writing systems is a persistent obstacle to the integration of multilingual geographic sources, whether modern gazetteers, medieval itineraries, or colonial-era surveys. Existing approaches depend on…

计算与语言 · 计算机科学 2026-03-31 Stephen Gadd

It is becoming common to archive research datasets that are not only large but also numerous. In addition, their corresponding metadata and the software required to analyse or display them need to be archived. Yet the manual curation of…

数字图书馆 · 计算机科学 2011-08-24 Daniel Lemire , Andre Vellino

The latest developments in digital have provided large data sets that can increasingly easily be accessed and used. These data sets often contain indirect localisation information, such as historical addresses. Historical geocoding is the…

数据库 · 计算机科学 2018-06-01 Rémi Cura , Bertrand Dumenieu , Nathalie Abadie , Benoit Costes , Julien Perret , Maurizio Gribaudi

We propose a global entity disambiguation (ED) model based on BERT. To capture global contextual information for ED, our model treats not only words but also entities as input tokens, and solves the task by sequentially resolving mentions…

计算与语言 · 计算机科学 2022-05-03 Ikuya Yamada , Koki Washio , Hiroyuki Shindo , Yuji Matsumoto

Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we developed a version…

计算与语言 · 计算机科学 2026-03-12 Masato Kikuchi , Masatsugu Ono , Toshioki Soga , Tetsu Tanabe , Tadachika Ozono
‹ 上一页 1 2 3 10 下一页 ›