English

Legal Evalutions and Challenges of Large Language Models

Computation and Language 2024-11-18 v1 Artificial Intelligence

Abstract

In this paper, we review legal testing methods based on Large Language Models (LLMs), using the OPENAI o1 model as a case study to evaluate the performance of large models in applying legal provisions. We compare current state-of-the-art LLMs, including open-source, closed-source, and legal-specific models trained specifically for the legal domain. Systematic tests are conducted on English and Chinese legal cases, and the results are analyzed in depth. Through systematic testing of legal cases from common law systems and China, this paper explores the strengths and weaknesses of LLMs in understanding and applying legal texts, reasoning through legal issues, and predicting judgments. The experimental results highlight both the potential and limitations of LLMs in legal applications, particularly in terms of challenges related to the interpretation of legal language and the accuracy of legal reasoning. Finally, the paper provides a comprehensive analysis of the advantages and disadvantages of various types of models, offering valuable insights and references for the future application of AI in the legal field.

Keywords

Cite

@article{arxiv.2411.10137,
  title  = {Legal Evalutions and Challenges of Large Language Models},
  author = {Jiaqi Wang and Huan Zhao and Zhenyuan Yang and Peng Shu and Junhao Chen and Haobo Sun and Ruixi Liang and Shixin Li and Pengcheng Shi and Longjun Ma and Zongjia Liu and Zhengliang Liu and Tianyang Zhong and Yutong Zhang and Chong Ma and Xin Zhang and Tuo Zhang and Tianli Ding and Yudan Ren and Tianming Liu and Xi Jiang and Shu Zhang},
  journal= {arXiv preprint arXiv:2411.10137},
  year   = {2024}
}