English

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

Multiagent Systems 2025-11-12 v1 Computation and Language

Abstract

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within complex tasks. We address this gap with a two-dimensional analysis of agent safety brittleness under the orthogonal pressures of intent concealment and task complexity. To enable this, we introduce OASIS (Orthogonal Agent Safety Inquiry Suite), a hierarchical benchmark with fine-grained annotations and a high-fidelity simulation sandbox. Our findings reveal two critical phenomena: safety alignment degrades sharply and predictably as intent becomes obscured, and a "Complexity Paradox" emerges, where agents seem safer on harder tasks only due to capability limitations. By releasing OASIS and its simulation environment, we provide a principled foundation for probing and strengthening agent safety in these overlooked dimensions.

Keywords

Cite

@article{arxiv.2511.08487,
  title  = {How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity},
  author = {Zihan Ma and Dongsheng Zhu and Shudong Liu and Taolin Zhang and Junnan Liu and Qingqiu Li and Minnan Luo and Songyang Zhang and Kai Chen},
  journal= {arXiv preprint arXiv:2511.08487},
  year   = {2025}
}
R2 v1 2026-07-01T07:32:33.709Z