The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Abstract
We introduce FACTS Grounding, an online leaderboard and associated benchmark that evaluates language models' ability to generate text that is factually accurate with respect to given context in the user prompt. In our benchmark, each prompt includes a user request and a full document, with a maximum length of 32k tokens, requiring long-form responses. The long-form responses are required to be fully grounded in the provided context document while fulfilling the user request. Models are evaluated using automated judge models in two phases: (1) responses are disqualified if they do not fulfill the user request; (2) they are judged as accurate if the response is fully grounded in the provided document. The automated judge models were comprehensively evaluated against a held-out test-set to pick the best prompt template, and the final factuality score is an aggregate of multiple judge models to mitigate evaluation bias. The FACTS Grounding leaderboard will be actively maintained over time, and contains both public and private splits to allow for external participation while guarding the integrity of the leaderboard. It can be found at https://www.kaggle.com/facts-leaderboard.
Keywords
Cite
@article{arxiv.2501.03200,
title = {The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input},
author = {Alon Jacovi and Andrew Wang and Chris Alberti and Connie Tao and Jon Lipovetz and Kate Olszewska and Lukas Haas and Michelle Liu and Nate Keating and Adam Bloniarz and Carl Saroufim and Corey Fry and Dror Marcus and Doron Kukliansky and Gaurav Singh Tomar and James Swirhun and Jinwei Xing and Lily Wang and Madhu Gurumurthy and Michael Aaron and Moran Ambar and Rachana Fellinger and Rui Wang and Zizhao Zhang and Sasha Goldshtein and Dipanjan Das},
journal= {arXiv preprint arXiv:2501.03200},
year = {2025}
}