English

TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness

Computation and Language 2024-05-08 v2

Abstract

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications. However, concerns have arisen regarding the trustworthiness of LLMs outputs, particularly in closed-book question-answering tasks, where non-experts may struggle to identify inaccuracies due to the absence of contextual or ground truth information. This paper introduces TrustScore, a framework based on the concept of Behavioral Consistency, which evaluates whether an LLMs response aligns with its intrinsic knowledge. Additionally, TrustScore can seamlessly integrate with fact-checking methods, which assesses alignment with external knowledge sources. The experimental results show that TrustScore achieves strong correlations with human judgments, surpassing existing reference-free metrics, and achieving results on par with reference-based metrics.

Keywords

Cite

@article{arxiv.2402.12545,
  title  = {TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness},
  author = {Danna Zheng and Danyang Liu and Mirella Lapata and Jeff Z. Pan},
  journal= {arXiv preprint arXiv:2402.12545},
  year   = {2024}
}
R2 v1 2026-06-28T14:53:47.372Z