English

Safe Vision Language Action Models via Barrier Enhanced Flow Matching

Robotics 2026-07-31 v1 Systems and Control

Abstract

This article presents a modular inference framework that integrates Flow Matching generative models with formal Control Barrier Function (CBF) safety guarantees. Unlike existing methods that apply external safety filters to a model's final output, our approach modifies the Flow Matching denoising process within the model to inherently generate safe trajectories. By employing a smooth Log-Sum-Exponential aggregate barrier, we enforce safety over entire action chunks. This aggregate barrier ensures a minimal increase in computational overhead and does not alter the semantic intent of the model. We show that, within the proposed framework, the 2-Wasserstein distance between the generated distribution and the target distribution remains bounded. Our method eliminates the need for safety-specific datasets or costly model retraining, providing a versatile solution for safe inference. We validate the approach on two robotic manipulation platforms and a 2D navigation benchmark, verifying that our framework achieves reliable safety without degrading the success rate of the model.

Cite

@article{arxiv.2607.29569,
  title  = {Safe Vision Language Action Models via Barrier Enhanced Flow Matching},
  author = {Kasra Sinaei and Hung-Chieh Wu and Donald Ebeigbe},
  journal= {arXiv preprint arXiv:2607.29569},
  year   = {2026}
}