中文
相关论文

相关论文: BullingerDB: A Dataset for Handwritten Text Recogn…

200 篇论文

The burgeoning integration of 3D medical imaging into healthcare has led to a substantial increase in the workload of medical professionals. To assist clinicians in their diagnostic processes and alleviate their workload, the development of…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yinda Chen , Che Liu , Xiaoyu Liu , Rossella Arcucci , Zhiwei Xiong

Modern BPE tokenizers often split calendar dates into meaningless fragments, e.g., 20250312 $\rightarrow$ 202, 503, 12, inflating token counts and obscuring the inherent structure needed for robust temporal reasoning. In this work, we (1)…

计算与语言 · 计算机科学 2025-09-25 Gagan Bhatia , Maxime Peyrard , Wei Zhao

The Covid-19 pandemic has caused a spur in the medical research literature. With new research advances in understanding the virus, there is a need for robust text mining tools which can process, extract and present answers from the…

信息检索 · 计算机科学 2021-08-04 Souvik Das , Sougata Saha , Rohini K. Srihari

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time.…

计算与语言 · 计算机科学 2026-02-13 Mariia Fedorova , Andrey Kutuzov , Khonzoda Umarova

This paper introduces a multi-level, multi-label text classification dataset comprising over 3000 documents. The dataset features literary and critical texts from 19th-century Ottoman Turkish and Russian. It is the first study to apply…

计算与语言 · 计算机科学 2024-07-23 Gokcen Gokceoglu , Devrim Cavusoglu , Emre Akbas , Özen Nergis Dolcerocca

Semantic annotation of long texts, such as novels, remains an open challenge in Natural Language Processing (NLP). This research investigates the problem of detecting person entities and assigning them unique identities, i.e., recognizing…

计算与语言 · 计算机科学 2021-10-05 Weronika Łajewska , Anna Wróblewska

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in document transcription,…

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliably distinguishing between human-authored and LLM-generated…

计算与语言 · 计算机科学 2024-12-18 Zhen Tao , Yanfang Chen , Dinghao Xi , Zhiyu Li , Wei Xu

Digital Forensics and Incident Response (DFIR) involves analyzing digital evidence to support legal investigations. Large Language Models (LLMs) offer new opportunities in DFIR tasks such as log analysis and memory forensics, but their…

密码学与安全 · 计算机科学 2025-05-27 Bilel Cherif , Tamas Bisztray , Richard A. Dubniczky , Aaesha Aldahmani , Saeed Alshehhi , Norbert Tihanyi

This paper deals with the task of practical and open source Handwritten Text Recognition (HTR) on German medieval manuscripts. We report on our efforts to construct mixed recognition models which can be applied out-of-the-box without any…

计算机视觉与模式识别 · 计算机科学 2022-01-20 Christian Reul , Stefan Tomasek , Florian Langhanki , Uwe Springmann

We present MultiTempBench, a multilingual temporal reasoning benchmark spanning three tasks, date arithmetic, time zone conversion, and temporal relation extraction across five languages (English, German, Chinese, Arabic, and Hausa) and…

计算与语言 · 计算机科学 2026-03-20 Gagan Bhatia , Ahmad Muhammad Isa , Maxime Peyrard , Wei Zhao

Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 documents are typically annotated. The number is constrained by time…

Unstructured data formats account for over 80% of the data currently stored, and extracting value from such formats remains a considerable challenge. In particular, current approaches for managing unstructured documents do not support…

With the widespread use of the internet, it has become increasingly crucial to extract specific information from vast amounts of academic articles efficiently. Data mining techniques are generally employed to solve this issue. However, data…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Jinghong Li , Koichi Ota , Wen Gu , Shinobu Hasegawa

Author name ambiguity in a digital library may affect the findings of research that mines authorship data of the library. This study evaluates author name disambiguation in DBLP, a widely used but insufficiently evaluated digital library…

数字图书馆 · 计算机科学 2018-07-31 Jinseok Kim

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce…

Most of the textual information available to us are temporally variable. In a world where information is dynamic, time-stamping them is a very important task. Documents are a good source of information and are used for many tasks like,…

计算与语言 · 计算机科学 2021-06-29 Swayambhu Nath Ray

This article presents a Bangla handwriting dataset named BanglaWriting that contains single-page handwritings of 260 individuals of different personalities and ages. Each page includes bounding-boxes that bounds each word, along with the…

计算机视觉与模式识别 · 计算机科学 2022-08-22 M. F. Mridha , Abu Quwsar Ohi , M. Ameer Ali , Mazedul Islam Emon , Muhammad Mohsin Kabir

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and…

To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The…