English

The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process

Computation and Language 2025-09-05 v1 Computers and Society Machine Learning

Abstract

Process Mining (PM), initially developed for industrial and business contexts, has recently been applied to social systems, including legal ones. However, PM's efficacy in the legal domain is limited by the accessibility and quality of datasets. We introduce ProLiFIC (Procedural Lawmaking Flow in Italian Chambers), a comprehensive event log of the Italian lawmaking process from 1987 to 2022. Created from unstructured data from the Normattiva portal and structured using large language models (LLMs), ProLiFIC aligns with recent efforts in integrating PM with LLMs. We exemplify preliminary analyses and propose ProLiFIC as a benchmark for legal PM, fostering new developments.

Keywords

Cite

@article{arxiv.2509.03528,
  title  = {The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process},
  author = {Matilde Contestabile and Chiara Ferrara and Alberto Giovannetti and Giovanni Parrillo and Andrea Vandin},
  journal= {arXiv preprint arXiv:2509.03528},
  year   = {2025}
}