English

IndianBailJudgments-1200: A Multi-Attribute Dataset for Legal NLP on Indian Bail Orders

Computation and Language 2025-07-04 v1 Artificial Intelligence Machine Learning

Abstract

Legal NLP remains underdeveloped in regions like India due to the scarcity of structured datasets. We introduce IndianBailJudgments-1200, a new benchmark dataset comprising 1200 Indian court judgments on bail decisions, annotated across 20+ attributes including bail outcome, IPC sections, crime type, and legal reasoning. Annotations were generated using a prompt-engineered GPT-4o pipeline and verified for consistency. This resource supports a wide range of legal NLP tasks such as outcome prediction, summarization, and fairness analysis, and is the first publicly available dataset focused specifically on Indian bail jurisprudence.

Keywords

Cite

@article{arxiv.2507.02506,
  title  = {IndianBailJudgments-1200: A Multi-Attribute Dataset for Legal NLP on Indian Bail Orders},
  author = {Sneha Deshmukh and Prathmesh Kamble},
  journal= {arXiv preprint arXiv:2507.02506},
  year   = {2025}
}

Comments

9 pages, 9 figures, 2 tables. Dataset available at Hugging Face and GitHub. Submitted to arXiv for open access