English

On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks

Computation and Language 2025-12-19 v3

Abstract

Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, including those involving human-AI interactions. In this paper, we present the first comprehensive study on the robustness of verbal confidence under adversarial attacks. We introduce attack frameworks targeting verbal confidence scores through both perturbation and jailbreak-based methods, and demonstrate that these attacks can significantly impair verbal confidence estimates and lead to frequent answer changes. We examine a variety of prompting strategies, model sizes, and application domains, revealing that current verbal confidence is vulnerable and that commonly used defence techniques are largely ineffective or counterproductive. Our findings underscore the need to design robust mechanisms for confidence expression in LLMs, as even subtle semantic-preserving modifications can lead to misleading confidence in responses.

Keywords

Cite

@article{arxiv.2507.06489,
  title  = {On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks},
  author = {Stephen Obadinma and Xiaodan Zhu},
  journal= {arXiv preprint arXiv:2507.06489},
  year   = {2025}
}

Comments

Published in NeurIPS 2025