English

Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models

Computer Vision and Pattern Recognition 2024-12-19 v1

Abstract

Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is closing in to near-perfect performance when considering hundreds of thousands of training samples. With the advent of large language models (LLMs), which are excellent few-shot learners, the question arises to what extent the document classification problem can be addressed with only a few training samples, or even none at all. In this paper, we investigate this question in the context of zero-shot prompting and few-shot model fine-tuning, with the aim of reducing the need for human-annotated training samples as much as possible.

Keywords

Cite

@article{arxiv.2412.13859,
  title  = {Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models},
  author = {Anna Scius-Bertrand and Michael Jungo and Lars Vögtlin and Jean-Marc Spat and Andreas Fischer},
  journal= {arXiv preprint arXiv:2412.13859},
  year   = {2024}
}

Comments

ICPR 2024

R2 v1 2026-06-28T20:40:29.418Z