English

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Artificial Intelligence 2026-05-28 v1

Abstract

Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnostic failure as citation laundering: a related source is presented as warrant for an over-strong claim. We introduce FORCEBENCH, a contrastive stress test for evidence-force calibration. Each item holds a cited passage fixed and pairs an evidence-calibrated claim with a localized force-raised variant across five operational axes: relation, modality, scope, temporal validity, and numeric specificity. A calibrated evaluator should score the evidence-calibrated claim higher. Headline experiments use a fixed, locality-filtered 198-pair evaluation set. A citation-presence sanity check is uninformative by design; token and entity overlap still violate monotonicity on 32.8--36.4% of pairs. Across four reported model judges, standard generic support prompting is insufficient for this force-calibration stress test (aggregate MVR 47.2%), while explicit warrant-strength prompting lowers MVR to 24.5% but remains imperfect. We release the benchmark, prompts, outputs, and plug-in pipeline so citation evaluators can report monotonicity violation rate and force sensitivity alongside conventional support metrics.

Keywords

Cite

@article{arxiv.2605.28044,
  title  = {Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG},
  author = {Pin Qian and Su Wang and Xiaoyuan Wang and Yihang Chen and Wenxuan Xu and Qiaolin Yu and Shuhuai Lin and Sipeng Zhang and Junxian You and Xinpeng Wei},
  journal= {arXiv preprint arXiv:2605.28044},
  year   = {2026}
}
R2 v1 2026-07-22T07:36:27.021Z