English

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

Cryptography and Security 2026-05-27 v1 Computation and Language

Abstract

In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.

Keywords

Cite

@article{arxiv.2605.27110,
  title  = {BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning},
  author = {Xuan Luo and Yue Wang and Geng Tu and Jing Li and Ruifeng Xu},
  journal= {arXiv preprint arXiv:2605.27110},
  year   = {2026}
}
R2 v1 2026-07-22T07:34:47.104Z