This paper conducts a comparative study on the performance of various machine learning (``ML'') approaches for classifying judgments into legal areas. Using a novel dataset of 6,227 Singapore Supreme Court judgments, we investigate how state-of-the-art NLP methods compare against traditional statistical models when applied to a legal corpus that comprised few but lengthy documents. All approaches tested, including topic model, word embedding, and language model-based classifiers, performed well with as little as a few hundred judgments. However, more work needs to be done to optimize state-of-the-art methods for the legal domain.
@article{arxiv.1904.06470,
title = {Legal Area Classification: A Comparative Study of Text Classifiers on Singapore Supreme Court Judgments},
author = {Jerrold Soh Tsin Howe and Lim How Khang and Ian Ernst Chai},
journal= {arXiv preprint arXiv:1904.06470},
year = {2019}
}
Comments
Accepted to the 1st Workshop on Natural Legal Language Processing (co-located with NAACL2019)