English

Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization

Machine Learning 2026-03-05 v1 Artificial Intelligence Cryptography and Security Multiagent Systems

Abstract

As Large Language Models (LLMs) transition into autonomous multi-agent ecosystems, robust minimax training becomes essential yet remains prone to instability when highly non-linear policies induce extreme local curvature in the inner maximization. Standard remedies that enforce global Jacobian bounds are overly conservative, suppressing sensitivity in all directions and inducing a large Price of Robustness. We introduce Adversarially-Aligned Jacobian Regularization (AAJR), a trajectory-aligned approach that controls sensitivity strictly along adversarial ascent directions. We prove that AAJR yields a strictly larger admissible policy class than global constraints under mild conditions, implying a weakly smaller approximation gap and reduced nominal performance degradation. Furthermore, we derive step-size conditions under which AAJR controls effective smoothness along optimization trajectories and ensures inner-loop stability. These results provide a structural theory for agentic robustness that decouples minimax stability from global expressivity restrictions.

Keywords

Cite

@article{arxiv.2603.04378,
  title  = {Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization},
  author = {Furkan Mumcu and Yasin Yilmaz},
  journal= {arXiv preprint arXiv:2603.04378},
  year   = {2026}
}
R2 v1 2026-07-01T11:03:35.211Z