English

A Survey of Document-Level Information Extraction

Computation and Language 2023-09-26 v1

Abstract

Document-level information extraction (IE) is a crucial task in natural language processing (NLP). This paper conducts a systematic review of recent document-level IE literature. In addition, we conduct a thorough error analysis with current state-of-the-art algorithms and identify their limitations as well as the remaining challenges for the task of document-level IE. According to our findings, labeling noises, entity coreference resolution, and lack of reasoning, severely affect the performance of document-level IE. The objective of this survey paper is to provide more insights and help NLP researchers to further enhance document-level IE performance.

Keywords

Cite

@article{arxiv.2309.13249,
  title  = {A Survey of Document-Level Information Extraction},
  author = {Hanwen Zheng and Sijia Wang and Lifu Huang},
  journal= {arXiv preprint arXiv:2309.13249},
  year   = {2023}
}
R2 v1 2026-06-28T12:30:09.673Z