中文
相关论文

相关论文: MedLatinEpi and MedLatinLit: Two Datasets for the …

200 篇论文

The Latin language has received attention from the computational linguistics research community, which has built, over the years, several valuable resources, ranging from detailed annotated corpora to sophisticated tools for linguistic…

计算与语言 · 计算机科学 2025-08-01 Alessandra Bassani , Beatrice Del Bo , Alfio Ferrara , Marta Mangini , Sergio Picascia , Ambra Stefanello

In the Middle Ages texts were learned by heart and spread using oral means of communication from generation to generation. Adaptation of the art of prose and poems allowed keeping particular descriptions and compositions characteristic for…

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual…

信息检索 · 计算机科学 2026-01-30 Julian Schelb , Michael Wittweiler , Marie Revellio , Barbara Feichtinger , Andreas Spitz

This paper describes the Quantitative Criticism Lab, a collaborative initiative between classicists, quantitative biologists, and computer scientists to apply ideas and methods drawn from the sciences to the study of literature. A core goal…

计算与语言 · 计算机科学 2023-06-22 Pramit Chaudhuri , Joseph P. Dexter

This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-shot authorship…

计算与语言 · 计算机科学 2024-10-15 Gleb Schmidt , Svetlana Gorovaia , Ivan P. Yamshchikov

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Maksym Voloshchuk , Bohdana Zarembovska , Mykola Kozlenko

In this work, we employ quantitative methods from the realm of statistics and machine learning to develop novel methodologies for author attribution and textual analysis. In particular, we develop techniques and software suitable for…

计算与语言 · 计算机科学 2014-05-06 James Brofos , Ajay Kannan , Rui Shu

The major hindrance in the study of earlier scientific literature is the availability of Latin translations into modern languages. This is particular true for the works of Euler who authored about 850 manuscripts and wrote a thousand…

历史与综述 · 数学 2024-04-22 Sylvio R. Bistafa

This work addresses critical challenges to academic integrity, including plagiarism, fabrication, and verification of authorship of educational content, by proposing a Natural Language Processing (NLP)-based framework for authenticating…

Large language models (LLMs) have gained popularity in various fields for their exceptional capability of generating human-like text. Their potential misuse has raised social concerns about plagiarism in academic contexts. However,…

人机交互 · 计算机科学 2023-06-02 Luoxuan Weng , Minfeng Zhu , Kam Kwai Wong , Shi Liu , Jiashun Sun , Hang Zhu , Dongming Han , Wei Chen

Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical…

The pharmacopeia used by physicians and lay people in medieval Europe has largely been dismissed as placebo or superstition. While we now recognise that some of the materia medica used by medieval physicians could have had useful biological…

定量方法 · 定量生物学 2018-07-20 Erin Connelly , Charo I. del Genio , Freya Harrison

Motivated by the sparsity of NLP resources for Eastern European languages, we present a broad index of existing Eastern European language resources (90+ datasets and 45+ models) published as a github repository open for updates from the…

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal…

计算与语言 · 计算机科学 2026-02-09 Yu Wu , Ke Shu , Jonas Fischer , Lidia Pivovarova , David Rosson , Eetu Mäkelä , Mikko Tolonen

It is well known that, within the Latin production of written text, peculiar metric schemes were followed not only in poetic compositions, but also in many prose works. Such metric patterns were based on so-called syllabic quantity, i.e.,…

计算与语言 · 计算机科学 2021-10-28 Silvia Corbara , Alejandro Moreo , Fabrizio Sebastiani

The computational analysis of poetry is limited by the scarcity of tools to automatically analyze and scan poems. In a multilingual settings, the problem is exacerbated as scansion and rhyme systems only exist for individual languages,…

计算与语言 · 计算机科学 2023-07-06 Javier de la Rosa , Álvaro Pérez Pozo , Salvador Ros , Elena González-Blanco

As language models become capable of processing increasingly long and complex texts, there has been growing interest in their application within computational literary studies. However, evaluating the usefulness of these models for such…

计算与语言 · 计算机科学 2026-01-21 Natasha Johnson , Amanda Bertsch , Maria-Emil Deal , Emma Strubell

One of the main drivers of the recent advances in authorship verification is the PAN large-scale authorship dataset. Despite generating significant progress in the field, inconsistent performance differences between the closed and open test…

计算与语言 · 计算机科学 2022-11-02 Florin Brad , Andrei Manolache , Elena Burceanu , Antonio Barbalau , Radu Ionescu , Marius Popescu

The research field concerned with the digital restoration of degraded written heritage lacks a quantitative metric for evaluating its results, which prevents the comparison of relevant methods on large datasets. Thus, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2021-02-22 Simon Brenner , Robert Sablatnig

With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predominantly stored in PDF formats and needs to be parsed into…

计算与语言 · 计算机科学 2025-02-04 Huawei Ji , Cheng Deng , Bo Xue , Zhouyang Jin , Jiaxin Ding , Xiaoying Gan , Luoyi Fu , Xinbing Wang , Chenghu Zhou
‹ 上一页 1 2 3 10 下一页 ›