English

Evalita-LLM: Benchmarking Large Language Models on Italian

Computation and Language 2025-02-05 v1

Abstract

We describe Evalita-LLM, a new benchmark designed to evaluate Large Language Models (LLMs) on Italian tasks. The distinguishing and innovative features of Evalita-LLM are the following: (i) all tasks are native Italian, avoiding issues of translating from Italian and potential cultural biases; (ii) in addition to well established multiple-choice tasks, the benchmark includes generative tasks, enabling more natural interaction with LLMs; (iii) all tasks are evaluated against multiple prompts, this way mitigating the model sensitivity to specific prompts and allowing a fairer and objective evaluation. We propose an iterative methodology, where candidate tasks and candidate prompts are validated against a set of LLMs used for development. We report experimental results from the benchmark's development phase, and provide performance statistics for several state-of-the-art LLMs.

Keywords

Cite

@article{arxiv.2502.02289,
  title  = {Evalita-LLM: Benchmarking Large Language Models on Italian},
  author = {Bernardo Magnini and Roberto Zanoli and Michele Resta and Martin Cimmino and Paolo Albano and Marco Madeddu and Viviana Patti},
  journal= {arXiv preprint arXiv:2502.02289},
  year   = {2025}
}

Comments

42 pages, 1 figure, 32 tables

R2 v1 2026-06-28T21:32:05.065Z