English

Evaluating the Reliability of Self-Explanations in Large Language Models

Computation and Language 2025-02-03 v2

Abstract

This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations - extractive and counterfactual - using three state-of-the-art LLMs (2B to 8B parameters) on two different classification tasks (objective and subjective). Our findings reveal, that, while these self-explanations can correlate with human judgement, they do not fully and accurately follow the model's decision process, indicating a gap between perceived and actual model reasoning. We show that this gap can be bridged because prompting LLMs for counterfactual explanations can produce faithful, informative, and easy-to-verify results. These counterfactuals offer a promising alternative to traditional explainability methods (e.g. SHAP, LIME), provided that prompts are tailored to specific tasks and checked for validity.

Keywords

Cite

@article{arxiv.2407.14487,
  title  = {Evaluating the Reliability of Self-Explanations in Large Language Models},
  author = {Korbinian Randl and John Pavlopoulos and Aron Henriksson and Tony Lindgren},
  journal= {arXiv preprint arXiv:2407.14487},
  year   = {2025}
}

Comments

Non peer-reviewed preprint. Presented at Discovery Science 2024. Peer-reviewed version published in the Springer Lecture Notes in Computer Science (vol 15243)

R2 v1 2026-06-28T17:47:38.334Z