English

ILSIC: Corpora for Identifying Indian Legal Statutes from Queries by Laypeople

Computation and Language 2026-02-03 v1

Abstract

Legal Statute Identification (LSI) for a given situation is one of the most fundamental tasks in Legal NLP. This task has traditionally been modeled using facts from court judgments as input queries, due to their abundance. However, in practical settings, the input queries are likely to be informal and asked by laypersons, or non-professionals. While a few laypeople LSI datasets exist, there has been little research to explore the differences between court and laypeople data for LSI. In this work, we create ILSIC, a corpus of laypeople queries covering 500+ statutes from Indian law. Additionally, the corpus also contains court case judgements to enable researchers to effectively compare between court and laypeople data for LSI. We conducted extensive experiments on our corpus, including benchmarking over the laypeople dataset using zero and few-shot inference, retrieval-augmented generation and supervised fine-tuning. We observe that models trained purely on court judgements are ineffective during test on laypeople queries, while transfer learning from court to laypeople data can be beneficial in certain scenarios. We also conducted fine-grained analyses of our results in terms of categories of queries and frequency of statutes.

Keywords

Cite

@article{arxiv.2602.00881,
  title  = {ILSIC: Corpora for Identifying Indian Legal Statutes from Queries by Laypeople},
  author = {Shounak Paul and Raghav Dogra and Pawan Goyal and Saptarshi Ghosh},
  journal= {arXiv preprint arXiv:2602.00881},
  year   = {2026}
}

Comments

9 Pages of Main, 1 page of Limitations and Ethics Statement, 11 Pages of Appendix, Accepted for Publication at EACL 2026 (Findings)