English

tabulapdf: An R Package to Extract Tables from PDF Documents

Information Retrieval 2024-09-24 v1 Digital Libraries

Abstract

tabulapdf is an R package that utilizes the Tabula Java library to import tables from PDF files directly into R. This tool can reduce time and effort in data extraction processes in fields like investigative journalism. It allows for automatic and manual table extraction, the latter facilitated through a Shiny interface, enabling manual areas selection with a computer mouse for data retrieval.

Cite

@article{arxiv.2409.14524,
  title  = {tabulapdf: An R Package to Extract Tables from PDF Documents},
  author = {Mauricio Vargas Sepúlveda and Thomas J. Leeper and Tom Paskhalis and Manuel Aristarán and Jeremy B. Merrill and Mike Tigas},
  journal= {arXiv preprint arXiv:2409.14524},
  year   = {2024}
}

Comments

10 pages, 1 figure

R2 v1 2026-06-28T18:52:59.890Z