Challenging the Abilities of Large Language Models in Italian: a Community Initiative
Abstract
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource -- the most comprehensive and diverse benchmark for Italian to date -- and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Keywords
Cite
@article{arxiv.2512.04759,
title = {Challenging the Abilities of Large Language Models in Italian: a Community Initiative},
author = {Malvina Nissim and Danilo Croce and Viviana Patti and Pierpaolo Basile and Giuseppe Attanasio and Elio Musacchio and Matteo Rinaldi and Federico Borazio and Maria Francis and Jacopo Gili and Daniel Scalena and Begoña Altuna and Ekhi Azurmendi and Valerio Basile and Luisa Bentivogli and Arianna Bisazza and Marianna Bolognesi and Dominique Brunato and Tommaso Caselli and Silvia Casola and Maria Cassese and Mauro Cettolo and Claudia Collacciani and Leonardo De Cosmo and Maria Pia Di Buono and Andrea Esuli and Julen Etxaniz and Chiara Ferrando and Alessia Fidelangeli and Simona Frenda and Achille Fusco and Marco Gaido and Andrea Galassi and Federico Galli and Luca Giordano and Mattia Goffetti and Itziar Gonzalez-Dios and Lorenzo Gregori and Giulia Grundler and Sandro Iannaccone and Chunyang Jiang and Moreno La Quatra and Francesca Lagioia and Soda Marem Lo and Marco Madeddu and Bernardo Magnini and Raffaele Manna and Fabio Mercorio and Paola Merlo and Arianna Muti and Vivi Nastase and Matteo Negri and Dario Onorati and Elena Palmieri and Sara Papi and Lucia Passaro and Giulia Pensa and Andrea Piergentili and Daniele Potertì and Giovanni Puccetti and Federico Ranaldi and Leonardo Ranaldi and Andrea Amelio Ravelli and Martina Rosola and Elena Sofia Ruzzetti and Giuseppe Samo and Andrea Santilli and Piera Santin and Gabriele Sarti and Giovanni Sartor and Beatrice Savoldi and Antonio Serino and Andrea Seveso and Lucia Siciliani and Paolo Torroni and Rossella Varvara and Andrea Zaninello and Asya Zanollo and Fabio Massimo Zanzotto and Kamyar Zeinalipour and Andrea Zugarini},
journal= {arXiv preprint arXiv:2512.04759},
year = {2025}
}