English

Data Augmented Pipeline for Legal Information Extraction and Reasoning

Computation and Language 2026-01-12 v1

Abstract

In this paper, we propose a pipeline leveraging Large Language Models (LLMs) for data augmentation in Information Extraction tasks within the legal domain. The proposed method is both simple and effective, significantly reducing the manual effort required for data annotation while enhancing the robustness of Information Extraction systems. Furthermore, the method is generalizable, making it applicable to various Natural Language Processing (NLP) tasks beyond the legal domain.

Keywords

Cite

@article{arxiv.2601.05609,
  title  = {Data Augmented Pipeline for Legal Information Extraction and Reasoning},
  author = {Nguyen Minh Phuong and Ha-Thanh Nguyen and May Myo Zin and Ken Satoh},
  journal= {arXiv preprint arXiv:2601.05609},
  year   = {2026}
}

Comments

Accepted in the Demonstration Track at ICAIL 2025

R2 v1 2026-07-01T08:57:28.187Z