中文

IL-PCSR:面向先例与法规检索的印度法律语料库

计算与语言 2025-11-04 v1 人工智能 信息检索 机器学习

摘要

识别/检索给定法律情形相关的法规和先例/先例是法律从业者常见的任务。迄今为止,研究人员分别独立地针对这两个任务, developing completely different datasets and models for each task; however, both retrieval tasks are inherently related, e.g., similar cases tend to cite similar statutes (due to similar factual situation). In this paper, we address this gap. We propose IL-PCR (Indian Legal corpus for Prior Case and Statute Retrieval), which is a unique corpus that provides a common testbed for developing models for both the tasks (Statute Retrieval and Precedent Retrieval) that can exploit the dependence between the two. We experiment extensively with several baseline models on the tasks, including lexical models, semantic models and ensemble based on GNNs. Further, to exploit the dependence between the two tasks, we develop an LLM-based re-ranking approach that gives the best performance.

关键词

引用

@article{arxiv.2511.00268,
  title  = {IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval},
  author = {Shounak Paul and Dhananjay Ghumare and Pawan Goyal and Saptarshi Ghosh and Ashutosh Modi},
  journal= {arXiv preprint arXiv:2511.00268},
  year   = {2025}
}

备注

Accepted at EMNLP 2025 (Main)