中文
相关论文

相关论文: Is text normalization relevant for classifying med…

200 篇论文

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic…

计算与语言 · 计算机科学 2025-02-26 Anton Ehrmanntraut

Although abbreviations are fairly common in handwritten sources, particularly in medieval and modern Western manuscripts, previous research dealing with computational approaches to their expansion is scarce. Yet abbreviations present…

计算与语言 · 计算机科学 2021-07-09 Jean-Baptiste Camps , Chahan Vidal-Gorène , Marguerite Vernet

Natural-language processing of historical documents is complicated by the abundance of variant spellings and lack of annotated data. A common approach is to normalize the spelling of historical words to modern forms. We explore the…

计算与语言 · 计算机科学 2016-10-26 Marcel Bollmann , Anders Søgaard

Identifying the production dates of historical manuscripts is one of the main goals for paleographers when studying ancient documents. Automatized methods can provide paleographers with objective tools to estimate dates more accurately.…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Lisa Koopmans , Maruf A. Dhali , Lambert Schomaker

This article focuses on the transcription of medieval manuscripts. Whereas problems of transcription have long interested medievalists, few workable options in the era of printed editions were available besides normalisation. The automation…

数字图书馆 · 计算机科学 2024-08-07 Estelle Guéville , David Joseph Wrisley

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Maksym Voloshchuk , Bohdana Zarembovska , Mykola Kozlenko

Diplomatics, the analysis of medieval charters, is a major field of research in which paleography is applied. Annotating data, if performed by laymen, needs validation and correction by experts. In this paper, we propose an effective and…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Anguelos Nicolaou , Daniel Luger , Franziska Decker , Nicolas Renet , Vincent Christlein , Georg Vogeler

This paper deals with the recognition and matching of text in both cartographic maps and ancient documents. The purpose of this work is to find similar text regions based on statistical and global features. A phase of normalization is done…

计算机视觉与模式识别 · 计算机科学 2013-08-30 Nizar Zaghden , Badreddine Khelifi , Adel M. Alimi , Remy Mullot

The Bavarian Academy of Sciences and Humanities aims to digitize its Medieval Latin Dictionary. This dictionary entails record cards referring to lemmas in medieval Latin, a low-resource language. A crucial step of the digitization process…

We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language. This is a fundamentally important routine to historians and digital humanities…

计算与语言 · 计算机科学 2022-01-25 Xutan Peng , Yi Zheng , Chenghua Lin , Advaith Siddharthan

This paper deals with the task of practical and open source Handwritten Text Recognition (HTR) on German medieval manuscripts. We report on our efforts to construct mixed recognition models which can be applied out-of-the-box without any…

计算机视觉与模式识别 · 计算机科学 2022-01-20 Christian Reul , Stefan Tomasek , Florian Langhanki , Uwe Springmann

In the realm of data privacy, the ability to effectively anonymise text is paramount. With the proliferation of deep learning and, in particular, transformer architectures, there is a burgeoning interest in leveraging these advanced models…

The Norman Conquest of 1066 C.E. brought profound transformations to England's administrative, societal, and linguistic practices. The DEEDS (Documents of Early England Data Set) database offers a unique opportunity to explore these changes…

计算与语言 · 计算机科学 2024-10-15 Yifan Liu , Gelila Tilahun , Xinxiang Gao , Qianfeng Wen , Michael Gervers

As more historical texts are digitized, there is interest in applying natural language processing tools to these archives. However, the performance of these tools is often unsatisfactory, due to language change and genre differences.…

计算与语言 · 计算机科学 2016-04-05 Yi Yang , Jacob Eisenstein

Several methods have been proposed for classifying long textual documents using Transformers. However, there is a lack of consensus on a benchmark to enable a fair comparison among different approaches. In this paper, we provide a…

计算与语言 · 计算机科学 2022-03-23 Hyunji Hayley Park , Yogarshi Vyas , Kashif Shah

Handwriting recognition is a key technology for accessing the content of old manuscripts, helping to preserve cultural heritage. Deep learning shows an impressive performance in solving this task. However, to achieve its full potential, it…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Michael Jungo , Lars Vögtlin , Atefeh Fakhari , Nathan Wegmann , Rolf Ingold , Andreas Fischer , Anna Scius-Bertrand

Short text classification is a crucial and challenging aspect of Natural Language Processing. For this reason, there are numerous highly specialized short text classifiers. However, in recent short text research, State of the Art (SOTA)…

计算与语言 · 计算机科学 2023-08-14 Fabian Karl , Ansgar Scherp

Transcription, annotation, digitization and/or visualization are common transformations that historical documents such as national records, birth/death registers, university records, letters or books undergo. Reasons for those…

人机交互 · 计算机科学 2020-09-07 Tomas Vancisin , Mary Orr , Uta Hinrichs

Text classification helps analyse texts for semantic meaning and relevance, by mapping the words against this hierarchy. An analysis of various types of texts is invaluable to understanding both their semantic meaning, as well as their…

机器学习 · 计算机科学 2022-11-16 Chaitanya Chadha , Vandit Gupta , Deepak Gupta , Ashish Khanna

Text classification is the task of assigning a document to a predefined class. However, it is expensive to acquire enough labeled documents or to label them. In this paper, we study the regularization methods' effects on various…

计算与语言 · 计算机科学 2024-03-05 Jongga Lee , Jaeseung Yim , Seohee Park , Changwon Lim
‹ 上一页 1 2 3 10 下一页 ›