English

FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation

Robotics 2026-07-09 v1

Abstract

We present FabriVLA, a lightweight Vision-Language-Action model for Precise Multi-Task Manipulation. FabriVLA combines an InternVL3.5 vision-language backbone with a flow-matching action head featuring gated self-attention across action tokens and shallow VLM layer fusion for enriched spatial context. The model is trained via single stage joint optimization from a pretrained VLM and randomly initialized action head. On the Meta-World MT50 benchmark spanning 50 diverse manipulation tasks, FabriVLA achieves a tier-average success rate of 90.0%, demonstrating that a compact VLA built on a 1B scale VLM can achieve strong performance without relying on multi billion parameter VLA backbones.

Cite

@article{arxiv.2607.08575,
  title  = {FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation},
  author = {Shiyuan Yang and Borong Zhang and Jizheng Zhang and Zhijia Tao and Junfei Guo and Donglai Ran and Xu Bian and Qingbiao Li},
  journal= {arXiv preprint arXiv:2607.08575},
  year   = {2026}
}