English
Related papers

Related papers: BullingerDB: A Dataset for Handwritten Text Recogn…

200 papers

The burgeoning integration of 3D medical imaging into healthcare has led to a substantial increase in the workload of medical professionals. To assist clinicians in their diagnostic processes and alleviate their workload, the development of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Yinda Chen , Che Liu , Xiaoyu Liu , Rossella Arcucci , Zhiwei Xiong

Modern BPE tokenizers often split calendar dates into meaningless fragments, e.g., 20250312 $\rightarrow$ 202, 503, 12, inflating token counts and obscuring the inherent structure needed for robust temporal reasoning. In this work, we (1)…

Computation and Language · Computer Science 2025-09-25 Gagan Bhatia , Maxime Peyrard , Wei Zhao

The Covid-19 pandemic has caused a spur in the medical research literature. With new research advances in understanding the virus, there is a need for robust text mining tools which can process, extract and present answers from the…

Information Retrieval · Computer Science 2021-08-04 Souvik Das , Sougata Saha , Rohini K. Srihari

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time.…

Computation and Language · Computer Science 2026-02-13 Mariia Fedorova , Andrey Kutuzov , Khonzoda Umarova

This paper introduces a multi-level, multi-label text classification dataset comprising over 3000 documents. The dataset features literary and critical texts from 19th-century Ottoman Turkish and Russian. It is the first study to apply…

Computation and Language · Computer Science 2024-07-23 Gokcen Gokceoglu , Devrim Cavusoglu , Emre Akbas , Özen Nergis Dolcerocca

Semantic annotation of long texts, such as novels, remains an open challenge in Natural Language Processing (NLP). This research investigates the problem of detecting person entities and assigning them unique identities, i.e., recognizing…

Computation and Language · Computer Science 2021-10-05 Weronika Łajewska , Anna Wróblewska

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in document transcription,…

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliably distinguishing between human-authored and LLM-generated…

Computation and Language · Computer Science 2024-12-18 Zhen Tao , Yanfang Chen , Dinghao Xi , Zhiyu Li , Wei Xu

Digital Forensics and Incident Response (DFIR) involves analyzing digital evidence to support legal investigations. Large Language Models (LLMs) offer new opportunities in DFIR tasks such as log analysis and memory forensics, but their…

Cryptography and Security · Computer Science 2025-05-27 Bilel Cherif , Tamas Bisztray , Richard A. Dubniczky , Aaesha Aldahmani , Saeed Alshehhi , Norbert Tihanyi

This paper deals with the task of practical and open source Handwritten Text Recognition (HTR) on German medieval manuscripts. We report on our efforts to construct mixed recognition models which can be applied out-of-the-box without any…

Computer Vision and Pattern Recognition · Computer Science 2022-01-20 Christian Reul , Stefan Tomasek , Florian Langhanki , Uwe Springmann

We present MultiTempBench, a multilingual temporal reasoning benchmark spanning three tasks, date arithmetic, time zone conversion, and temporal relation extraction across five languages (English, German, Chinese, Arabic, and Hausa) and…

Computation and Language · Computer Science 2026-03-20 Gagan Bhatia , Ahmad Muhammad Isa , Maxime Peyrard , Wei Zhao

Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 documents are typically annotated. The number is constrained by time…

Computation and Language · Computer Science 2026-01-23 Jaya Chaturvedi , Saniya Deshpande , Chenkai Ma , Robert Cobb , Angus Roberts , Robert Stewart , Daniel Stahl , Diana Shamsutdinova

Unstructured data formats account for over 80% of the data currently stored, and extracting value from such formats remains a considerable challenge. In particular, current approaches for managing unstructured documents do not support…

With the widespread use of the internet, it has become increasingly crucial to extract specific information from vast amounts of academic articles efficiently. Data mining techniques are generally employed to solve this issue. However, data…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Jinghong Li , Koichi Ota , Wen Gu , Shinobu Hasegawa

Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library…

Digital Libraries · Computer Science 2018-07-31 Jinseok Kim

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce…

Computation and Language · Computer Science 2026-03-03 Jason Lucas , Matt Murtagh-White , Adaku Uchendu , Ali Al-Lawati , Michiharu Yamashita , Dominik Macko , Ivan Srba , Robert Moro , Dongwon Lee

Most of the textual information available to us are temporally variable. In a world where information is dynamic, time-stamping them is a very important task. Documents are a good source of information and are used for many tasks like,…

Computation and Language · Computer Science 2021-06-29 Swayambhu Nath Ray

This article presents a Bangla handwriting dataset named BanglaWriting that contains single-page handwritings of 260 individuals of different personalities and ages. Each page includes bounding-boxes that bounds each word, along with the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 M. F. Mridha , Abu Quwsar Ohi , M. Ameer Ali , Mazedul Islam Emon , Muhammad Mohsin Kabir

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and…

To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The…

Information Retrieval · Computer Science 2025-11-19 Dawn Lawrie , James Mayfield , Eugene Yang , Andrew Yates , Sean MacAvaney , Ronak Pradeep , Scott Miller , Paul McNamee , Luca Soldani