中文
相关论文

相关论文: HORAE: an annotated dataset of books of hours

200 篇论文

Segmentation of Arabic manuscripts into lines of text and words is an important step to make recognition systems more efficient and accurate. The problem of segmentation into text lines is solved since there are carefully annotated dataset…

计算与语言 · 计算机科学 2023-12-14 Hakim Bouchal , Ahror Belaid

Calligraphy is an essential part of the Arabic heritage and culture. It has been used in the past for the decoration of houses and mosques. Usually, such calligraphy is designed manually by experts with aesthetic insights. In the past few…

计算与语言 · 计算机科学 2021-06-28 Zaid Alyafeai , Maged S. Al-shaibani , Mustafa Ghaleb , Yousif Ahmed Al-Wajih

We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Mehreen Saeed , Adrian Chan , Anupam Mijar , Joseph Moukarzel , Georges Habchi , Carlos Younes , Amin Elias , Chau-Wai Wong , Akram Khater

Handwritten text recognition for historical documents is an important task but it remains difficult due to a lack of sufficient training data in combination with a large variability of writing styles and degradation of historical documents.…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Christian M. Dahl , Torben S. D. Johansen , Emil N. Sørensen , Christian E. Westermann , Simon F. Wittrock

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

计算与语言 · 计算机科学 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

Research units in archaeology often manage large and precious archives containing various documents, including reports on fieldwork, scholarly studies and reference books. These archives are of course invaluable, recording decades of work,…

数字图书馆 · 计算机科学 2015-07-09 Frédérique Mélanie-Becquet , Johan Ferguth , Katherine Gruel , Thierry Poibeau

There is a large volume of late antique and medieval Hebrew texts. They represent a crucial linguistic and cultural bridge between Biblical and modern Hebrew. Poetry is prominent in these texts and one of its main haracteristics is the…

计算与语言 · 计算机科学 2025-03-04 Michael Toker , Oren Mishali , Ophir Münz-Manor , Benny Kimelfeld , Yonatan Belinkov

The Latin language has received attention from the computational linguistics research community, which has built, over the years, several valuable resources, ranging from detailed annotated corpora to sophisticated tools for linguistic…

计算与语言 · 计算机科学 2025-08-01 Alessandra Bassani , Beatrice Del Bo , Alfio Ferrara , Marta Mangini , Sergio Picascia , Ambra Stefanello

Archives are an important source of study for various scholars. Digitization and the web have made archives more accessible and led to the development of several time-aware exploratory search systems. However these systems have been…

信息检索 · 计算机科学 2018-10-26 Jaspreet Singh , Wolfgang Nejdl , Avishek Anand

We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Martin Kišš , Michal Hradiš , Martina Dvořáková , Václav Jiroušek , Filip Kersch

We report on a language resource consisting of 2000 annotated bibliography entries, which is being analyzed as part of our research on indicative document summarization. We show how annotated bibliographies cover certain aspects of…

计算与语言 · 计算机科学 2007-05-23 Min-Yen Kan , Judith L. Klavans , Kathleen R. McKeown

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus of 3.6M sentences…

计算与语言 · 计算机科学 2023-06-08 Vít Novotný , Kristýna Luger , Michal Štefánik , Tereza Vrabcová , Aleš Horák

This paper deals with the task of practical and open source Handwritten Text Recognition (HTR) on German medieval manuscripts. We report on our efforts to construct mixed recognition models which can be applied out-of-the-box without any…

计算机视觉与模式识别 · 计算机科学 2022-01-20 Christian Reul , Stefan Tomasek , Florian Langhanki , Uwe Springmann

In the Middle Ages texts were learned by heart and spread using oral means of communication from generation to generation. Adaptation of the art of prose and poems allowed keeping particular descriptions and compositions characteristic for…

This article focuses on the transcription of medieval manuscripts. Whereas problems of transcription have long interested medievalists, few workable options in the era of printed editions were available besides normalisation. The automation…

数字图书馆 · 计算机科学 2024-08-07 Estelle Guéville , David Joseph Wrisley

Understanding human behavior is key for robots and intelligent systems that share a space with people. Accordingly, research that enables such systems to perceive, track, learn and predict human behavior as well as to plan and interact with…

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including…

计算机视觉与模式识别 · 计算机科学 2020-09-14 James P. Philips , Nasseh Tabrizi

We present a free and open-source tool for creating web-based surveys that include text annotation tasks. Existing tools offer either text annotation or survey functionality but not both. Combining the two input types is particularly…

计算与语言 · 计算机科学 2021-12-20 Timo Spinde , Kanishka Sinha , Norman Meuschke , Bela Gipp

Historical materials are abundant. Yet, piecing together how human knowledge has evolved and spread both diachronically and synchronically remains a challenge that can so far only be very selectively addressed. The vast volume of materials…

This piece plays with the idea of the Computocene: an era defined not merely by the ubiquity of computers, but by their deepening role in how we observe, interpret, and make sense of the world. Rather than emphasizing automation, speed,…

计算机与社会 · 计算机科学 2025-05-29 Simone Severini
‹ 上一页 1 2 3 10 下一页 ›