English
Related papers

Related papers: Is text normalization relevant for classifying med…

200 papers

Accessibility to historical documents is mostly limited to scholars. This is due to the language barrier inherent in human language and the linguistic properties of these documents. Given a historical document, modernization aims to…

Computation and Language · Computer Science 2020-03-05 Miguel Domingo , Francisco Casacuberta

One advantage of neural ranking models is that they are meant to generalise well in situations of synonymity i.e. where two words have similar or identical meanings. In this paper, we investigate and quantify how well various ranking models…

Information Retrieval · Computer Science 2023-08-02 Andreas Chari , Sean MacAvaney , Iadh Ounis

The automated categorization (or classification) of texts into predefined categories has witnessed a booming interest in the last ten years, due to the increased availability of documents in digital form and the ensuing need to organize…

Information Retrieval · Computer Science 2021-09-21 Fabrizio Sebastiani

While transformer-based models achieve strong performance on text classification, we explore whether masking input tokens can further enhance their effectiveness. We propose token masking regularization, a simple yet theoretically motivated…

Computation and Language · Computer Science 2025-05-20 Xianglong Xu , John Bowen , Rojin Taheri

Alterations in historical manuscripts such as letters represent a promising field of research. On the one hand, they help understand the construction of text. On the other hand, topics that are being considered sensitive at the time of the…

Machine Learning · Computer Science 2020-11-05 David Lassner , Anne Baillot , Sergej Dogadov , Klaus-Robert Müller , Shinichi Nakajima

The goal of text ranking is to generate an ordered list of texts retrieved from a corpus in response to a query. Although the most common formulation of text ranking is search, instances of the task can also be found in many natural…

Information Retrieval · Computer Science 2021-08-20 Jimmy Lin , Rodrigo Nogueira , Andrew Yates

We investigate the pertinence of methods from algebraic topology for text data analysis. These methods enable the development of mathematically-principled isometric-invariant mappings from a set of vectors to a document embedding, which is…

Computation and Language · Computer Science 2017-06-01 Paul Michel , Abhilasha Ravichander , Shruti Rijhwani

In this paper we consider two sequence tagging tasks for medieval Latin: part-of-speech tagging and lemmatization. These are both basic, yet foundational preprocessing steps in applications such as text re-use detection. Nevertheless, they…

Computation and Language · Computer Science 2023-06-22 Mike Kestemont , Jeroen De Gussem

In recent years, large pre-trained transformers have led to substantial gains in performance over traditional retrieval models and feedback approaches. However, these results are primarily based on the MS Marco/TREC Deep Learning Track…

Information Retrieval · Computer Science 2022-04-18 David Rau , Jaap Kamps

Violence descriptions in literature offer valuable insights for a wide range of research in the humanities. For historians, depictions of violence are of special interest for analyzing the societal dynamics surrounding large wars and…

Computation and Language · Computer Science 2025-03-12 Alhassan Abdelhalim , Michaela Regneri

Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as automated assessors,…

Information Retrieval · Computer Science 2025-12-08 Samaneh Mohtadi , Kevin Roitero , Stefano Mizzaro , Gianluca Demartini

Maps are an important source of information in archaeology and other sciences. Users want to search for historical maps to determine recorded history of the political geography of regions at different eras, to find out where exactly…

Digital Libraries · Computer Science 2009-01-27 Qingzhao Tan , Prasenjit Mitra , C. Lee Giles

Recent improvements in the quality of the generations by large language models have spurred research into identifying machine-generated text. Such work often presents high-performing detectors. However, humans and machines can produce text…

Computation and Language · Computer Science 2024-12-13 Jad Doughman , Osama Mohammed Afzal , Hawau Olamide Toyin , Shady Shehata , Preslav Nakov , Zeerak Talat

Text classification has seen an increased use in both academic and industry settings. Though rule based methods have been fairly successful, supervised machine learning has been shown to be most successful for most languages, where most…

Computation and Language · Computer Science 2021-04-26 Deniz Kavi

The presence of specific linguistic signals particular to a certain sub-group can become highly salient to language models during training. In automated decision-making settings, this may lead to biased outcomes when models rely on cues…

Computation and Language · Computer Science 2025-09-05 Charmaine Barker , Dimitar Kazakov

The Normalization transformation plays a key role in the compilation of Diderot programs. The transformations are complicated and it would be easy for a bug to go undetected. To increase our confidence in normalization part of the compiler…

Programming Languages · Computer Science 2017-05-25 Charisee Chiw , John Reppy

The vulnerability of models to data aberrations and adversarial attacks influences their ability to demarcate distinct class boundaries efficiently. The network's confidence and uncertainty play a pivotal role in weight adjustments and the…

Machine Learning · Computer Science 2020-12-15 Utkarsh Uppal , Bharat Giddwani

Scene text recognition has made significant progress in recent years and has become an important part of the work-flow. The widespread use of mobile devices opens up wide possibilities for using OCR technologies in everyday life. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-03 Nathan Zachary , Gerald Carl , Russell Elijah , Hessi Roma , Robert Leer , James Amelia

This study introduces the eFontes models for automatic linguistic annotation of Medieval Latin texts, focusing on lemmatization, part-of-speech tagging, and morphological feature determination. Using the Transformers library, these models…

Computation and Language · Computer Science 2024-07-02 Krzysztof Nowak , Jędrzej Ziębura , Krzysztof Wróbel , Aleksander Smywiński-Pohl

Training data for text classification is often limited in practice, especially for applications with many output classes or involving many related classification problems. This means classifiers must generalize from limited evidence, but…

Computation and Language · Computer Science 2020-05-19 Abhijit Mahabal , Jason Baldridge , Burcu Karagol Ayan , Vincent Perot , Dan Roth