English

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Computation and Language 2026-05-12 v1 Artificial Intelligence

Abstract

Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competing claims, applying doctrine to facts, and rendering reasoned decisions -- is arguably as fundamental to a well-functioning legal system as advocacy itself. We introduce Magis-Bench, a benchmark for evaluating LLMs on magistrate-level writing tasks derived from recent Brazilian competitive examinations for judicial positions. Magis-Bench comprises 74 questions from eight examinations conducted between 2023 and 2025, including discursive legal analysis questions with multi-turn structure and practical exercises requiring the composition of complete civil and criminal judicial sentences. We evaluate 23 state-of-the-art LLMs using an LLM-as-a-judge methodology with four independent frontier models as evaluators. Our results show strong inter-judge agreement (Kendall's W=0.984W = 0.984; pairwise Kendall's τ0.897\tau \ge 0.897), with Google's Gemini-3-Pro-Preview achieving the highest average score (6.97/10), followed by Gemini-3-Flash-Preview (6.67) and Claude-4.5-Opus (6.46). Even the best-performing models score below 70\% of the maximum, indicating that judicial-level legal reasoning and writing remain challenging for current LLMs. We release the complete benchmark, model outputs, and evaluation code to support further research on legal AI capabilities.

Keywords

Cite

@article{arxiv.2605.08437,
  title  = {Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks},
  author = {Ramon Pires and Thales Sales Almeida and Celio Larcher Junior and Giovana Bonás and Hugo Abonizio and Marcos Piau and Roseval Malaquias Junior and Thiago Laitz and Rodrigo Nogueira},
  journal= {arXiv preprint arXiv:2605.08437},
  year   = {2026}
}