中文
相关论文

相关论文: LiMe: a Latin Corpus of Late Medieval Criminal Sen…

200 篇论文

In the Middle Ages texts were learned by heart and spread using oral means of communication from generation to generation. Adaptation of the art of prose and poems allowed keeping particular descriptions and compositions characteristic for…

We present and make available MedLatinEpi and MedLatinLit, two datasets of medieval Latin texts to be used in research on computational authorship analysis. MedLatinEpi and MedLatinLit consist of 294 and 30 curated texts, respectively,…

计算与语言 · 计算机科学 2021-09-22 Silvia Corbara , Alejandro Moreo , Fabrizio Sebastiani , Mirko Tavoni

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual…

信息检索 · 计算机科学 2026-01-30 Julian Schelb , Michael Wittweiler , Marie Revellio , Barbara Feichtinger , Andreas Spitz

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances…

计算与语言 · 计算机科学 2020-09-22 David Bamman , Patrick J. Burns

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal…

计算与语言 · 计算机科学 2026-02-09 Yu Wu , Ke Shu , Jonas Fischer , Lidia Pivovarova , David Rosson , Eetu Mäkelä , Mikko Tolonen

Large language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination. In this work, our aim is to allow LLMs to generate text with citations, improving their factual…

计算与语言 · 计算机科学 2023-11-01 Tianyu Gao , Howard Yen , Jiatong Yu , Danqi Chen

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

计算与语言 · 计算机科学 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Maksym Voloshchuk , Bohdana Zarembovska , Mykola Kozlenko

In this paper we consider two sequence tagging tasks for medieval Latin: part-of-speech tagging and lemmatization. These are both basic, yet foundational preprocessing steps in applications such as text re-use detection. Nevertheless, they…

计算与语言 · 计算机科学 2023-06-22 Mike Kestemont , Jeroen De Gussem

The Bavarian Academy of Sciences and Humanities aims to digitize its Medieval Latin Dictionary. This dictionary entails record cards referring to lemmas in medieval Latin, a low-resource language. A crucial step of the digitization process…

Large language models have gained tremendous popularity in domains such as e-commerce, finance, healthcare, and education. Fine-tuning is a common approach to customize an LLM on a domain-specific dataset for a desired downstream task. In…

计算与语言 · 计算机科学 2024-06-11 Shraboni Sarker , Ahmad Tamim Hamad , Hulayyil Alshammari , Viviana Grieco , Praveen Rao

In this paper, we aim at the application of Natural Language Processing (NLP) techniques to historical research endeavors, particularly addressing the study of religious invectives in the context of the Protestant Reformation in Tudor…

计算与语言 · 计算机科学 2025-09-29 Sophie Spliethoff , Sanne Hoeken , Silke Schwandt , Sina Zarrieß , Özge Alaçam

The record of the beginning of the most widespread legal system in the world is contained in millions of pages of handwritten text. Most of the records of the first centuries of the Anglo-American legal system are hand-written in a highly…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Michael Zhang , Elise Wang , Charlotte Whatley , Seth Strickland , Dylan Bannon

In this article we present the Frankfurt Latin Lexicon (FLL), a lexical resource for Medieval Latin that is used both for the lemmatization of Latin texts and for the post-editing of lemmatizations. We describe recent advances in the…

This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-shot authorship…

计算与语言 · 计算机科学 2024-10-15 Gleb Schmidt , Svetlana Gorovaia , Ivan P. Yamshchikov

The IMPACT-es diachronic corpus of historical Spanish compiles over one hundred books --containing approximately 8 million words-- in addition to a complementary lexicon which links more than 10 thousand lemmas with attestations of the…

计算与语言 · 计算机科学 2013-07-01 Felipe Sánchez-Martínez , Isabel Martínez-Sempere , Xavier Ivars-Ribes , Rafael C. Carrasco

The study investigates the efficacy of pre-trained language models (PLMs) in analyzing argumentative moves in a longitudinal learner corpus. Prior studies on argumentative moves often rely on qualitative analysis and manual coding, limiting…

计算与语言 · 计算机科学 2025-06-04 Wenjuan Qin , Weiran Wang , Yuming Yang , Tao Gui

This study introduces the eFontes models for automatic linguistic annotation of Medieval Latin texts, focusing on lemmatization, part-of-speech tagging, and morphological feature determination. Using the Transformers library, these models…

计算与语言 · 计算机科学 2024-07-02 Krzysztof Nowak , Jędrzej Ziębura , Krzysztof Wróbel , Aleksander Smywiński-Pohl

We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper…

In this study, we propose to evaluate the use of deep learning methods for semantic classification at the sentence level to accelerate the process of corpus building in the field of humanities and linguistics, a traditional and…

计算与语言 · 计算机科学 2024-03-27 Thibault Clérice
‹ 上一页 1 2 3 10 下一页 ›