English

Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System

Cryptography and Security 2025-11-18 v2 Computation and Language

Abstract

State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents \textsc{Fact2Fiction}, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9\%--21.2\% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.

Keywords

Cite

@article{arxiv.2508.06059,
  title  = {Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System},
  author = {Haorui He and Yupeng Li and Bin Benjamin Zhu and Dacheng Wen and Reynold Cheng and Francis C. M. Lau},
  journal= {arXiv preprint arXiv:2508.06059},
  year   = {2025}
}

Comments

Accepted by AAAI 2026 (Oral). Code available at: https://trustworthycomp.github.io/Fact2Fiction/

R2 v1 2026-07-01T04:40:29.081Z