English
Related papers

Related papers: PubMed-OCR: PMC Open Access OCR Annotations

200 papers

OpenCitations is an infrastructure organization for open scholarship dedicated to the publication of open citation data as Linked Open Data using Semantic Web technologies, thereby providing a disruptive alternative to traditional…

Digital Libraries · Computer Science 2020-02-24 Silvio Peroni , David Shotton

Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. However, real-world documents often feature complex layouts…

Since the dawn of the computing era, information has been represented digitally so that it can be processed by electronic computers. Paper books and documents were abundant and widely being published at that time; and hence, there was a…

Computation and Language · Computer Science 2012-04-03 Youssef Bassil , Mohammad Alwani

Retrieving accurate details from documents is a crucial task, especially when handling a combination of scanned images and native digital formats. This document presents a combined framework for text extraction that merges Optical Character…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Rasha Sinha , Rekha B S

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean…

Computation and Language · Computer Science 2020-06-09 Hillel Taub-Tabib , Micah Shlain , Shoval Sadde , Dan Lahav , Matan Eyal , Yaara Cohen , Yoav Goldberg

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other…

Computation and Language · Computer Science 2023-08-25 Melissa Dell , Jacob Carlson , Tom Bryan , Emily Silcock , Abhishek Arora , Zejiang Shen , Luca D'Amico-Wong , Quan Le , Pablo Querubin , Leander Heldring

There has been rapid growth in biomedical literature, yet capturing the heterogeneity of the bibliographic information of these articles remains relatively understudied. Although graph mining research via heterogeneous graph neural networks…

Machine Learning · Computer Science 2023-08-28 Eric W Lee , Joyce C Ho

Wikipedia is an essential component of the open science ecosystem, yet it is poorly integrated with academic open science initiatives. Wikipedia Citations is a project that focuses on extracting and releasing comprehensive datasets of…

Digital Libraries · Computer Science 2024-06-28 Natallia Kokash , Giovanni Colavizza

PubTator 3.0 (https://www.ncbi.nlm.nih.gov/research/pubtator3/) is a biomedical literature resource using state-of-the-art AI techniques to offer semantic and relation searches for key concepts like proteins, genetic variants, diseases, and…

Computation and Language · Computer Science 2024-01-23 Chih-Hsuan Wei , Alexis Allot , Po-Ting Lai , Robert Leaman , Shubo Tian , Ling Luo , Qiao Jin , Zhizheng Wang , Qingyu Chen , Zhiyong Lu

We present a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library. In the case reports, we annotate cases, conditions, findings, factors and negation modifiers.…

Computation and Language · Computer Science 2020-03-31 Sarah Schulz , Jurica Ševa , Samuel Rodriguez , Malte Ostendorff , Georg Rehm

Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers,…

Digital Libraries · Computer Science 2025-09-17 Fitsum Sileshi Beyene , Christopher L. Dancy

Financial documents are essential sources of information for regulators, auditors, and financial institutions, particularly for assessing the wealth and compliance of Small and Medium-sized Businesses. However, SMB documents are often…

Information Retrieval · Computer Science 2025-10-28 Yichao Jin , Yushuo Wang , Qishuai Zhong , Kent Chiu Jin-Chun , Kenneth Zhu Ke , Donald MacDonald

Medical Subject Heading (MeSH) indexing refers to the problem of assigning a given biomedical document with the most relevant labels from an extremely large set of MeSH terms. Currently, the vast number of biomedical articles in the PubMed…

Computation and Language · Computer Science 2022-04-29 Xindi Wang , Robert E. Mercer , Frank Rudzicz

Optical Character Recognition (OCR) is a critical but error-prone stage in digital humanities text pipelines. While OCR correction improves usability for downstream NLP tasks, common workflows often overwrite intermediate decisions,…

Human-Computer Interaction · Computer Science 2026-05-07 Haoze Guo , Ziqi Wei

This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the era of large language models, PaddleOCR 3.0 presents three…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Cheng Cui , Ting Sun , Manhui Lin , Tingquan Gao , Yubo Zhang , Jiaxuan Liu , Xueqing Wang , Zelun Zhang , Changda Zhou , Hongen Liu , Yue Zhang , Wenyu Lv , Kui Huang , Yichao Zhang , Jing Zhang , Jun Zhang , Yi Liu , Dianhai Yu , Yanjun Ma

The need for robust and diverse data sets to train clinical large language models (cLLMs) is critical given that currently available public repositories often prove too limited in size or scope for comprehensive medical use. While resources…

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is crucial for assessing…

Computation and Language · Computer Science 2026-02-04 Deniz Yılmaz , Evren Ayberk Munis , Çağrı Toraman , Süha Kağan Köse , Burak Aktaş , Mehmet Can Baytekin , Bilge Kaan Görür

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close…

Digital Libraries · Computer Science 2020-11-24 Harshdeep Singh , Robert West , Giovanni Colavizza

Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows. These documents are typically non machine readable, noisy, and visually heterogeneous.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Yuxuan Han , Yuanxing Zhang , Yushuo Wang , Yichao Jin , Kenneth Zhu Ke , Jingyuan Zhao