English

Quantifying True Robustness: Synonymity-Weighted Similarity for Trustworthy XAI Evaluation

Machine Learning 2025-12-30 v2 Artificial Intelligence Computation and Language

Abstract

Adversarial attacks challenge the reliability of Explainable AI (XAI) by altering explanations while the model's output remains unchanged. The success of these attacks on text-based XAI is often judged using standard information retrieval metrics. We argue these measures are poorly suited in the evaluation of trustworthiness, as they treat all word perturbations equally while ignoring synonymity, which can misrepresent an attack's true impact. To address this, we apply synonymity weighting, a method that amends these measures by incorporating the semantic similarity of perturbed words. This produces more accurate vulnerability assessments and provides an important tool for assessing the robustness of AI systems. Our approach prevents the overestimation of attack success, leading to a more faithful understanding of an XAI system's true resilience against adversarial manipulation.

Keywords

Cite

@article{arxiv.2501.01516,
  title  = {Quantifying True Robustness: Synonymity-Weighted Similarity for Trustworthy XAI Evaluation},
  author = {Christopher Burger},
  journal= {arXiv preprint arXiv:2501.01516},
  year   = {2025}
}

Comments

10 pages, 2 figures, 6 tables. Changes to title, abstract and minor edits to the content as a result of acceptance to the 59th Hawaii International Conference on System Sciences