The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
Abstract
We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually accurate text across diverse scenarios. The suite provides a holistic measure of factuality by aggregating the performance of models on four distinct sub-leaderboards: (1) FACTS Multimodal, which measures the factuality of responses to image-based questions; (2) FACTS Parametric, which assesses models' world knowledge by answering closed-book factoid questions from internal parameters; (3) FACTS Search, which evaluates factuality in information-seeking scenarios, where the model must use a search API; and (4) FACTS Grounding (v2), which evaluates whether long-form responses are grounded in provided documents, featuring significantly improved judge models. Each sub-leaderboard employs automated judge models to score model responses, and the final suite score is an average of the four components, designed to provide a robust and balanced assessment of a model's overall factuality. The FACTS Leaderboard Suite will be actively maintained, containing both public and private splits to allow for external participation while guarding its integrity. It can be found at https://www.kaggle.com/benchmarks/google/facts .
Cite
@article{arxiv.2512.10791,
title = {The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality},
author = {Aileen Cheng and Alon Jacovi and Amir Globerson and Ben Golan and Charles Kwong and Chris Alberti and Connie Tao and Eyal Ben-David and Gaurav Singh Tomar and Lukas Haas and Yonatan Bitton and Adam Bloniarz and Aijun Bai and Andrew Wang and Anfal Siddiqui and Arturo Bajuelos Castillo and Aviel Atias and Chang Liu and Corey Fry and Daniel Balle and Deepanway Ghosal and Doron Kukliansky and Dror Marcus and Elena Gribovskaya and Eran Ofek and Honglei Zhuang and Itay Laish and Jan Ackermann and Lily Wang and Meg Risdal and Megan Barnes and Michael Fink and Mohamed Amin and Moran Ambar and Natan Potikha and Nikita Gupta and Nitzan Katz and Noam Velan and Ofir Roval and Ori Ram and Polina Zablotskaia and Prathamesh Bang and Priyanka Agrawal and Rakesh Ghiya and Sanjay Ganapathy and Simon Baumgartner and Sofia Erell and Sushant Prakash and Thibault Sellam and Vikram Rao and Xuanhui Wang and Yaroslav Akulov and Yulong Yang and Zhen Yang and Zhixin Lai and Zhongru Wu and Anca Dragan and Avinatan Hassidim and Fernando Pereira and Slav Petrov and Srinivasan Venkatachary and Tulsee Doshi and Yossi Matias and Sasha Goldshtein and Dipanjan Das},
journal= {arXiv preprint arXiv:2512.10791},
year = {2025}
}