English

RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models

Distributed, Parallel, and Cluster Computing 2026-03-13 v2 Robotics

Abstract

Vision Language Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) inference offers an effective fix by easing edge-device computing pressure to meet real-time needs. However, existing ECC frameworks are suboptimal for VLA models due to two challenges: (1) Mainstream environment-oriented edge-cloud partitioning methods are susceptible to interference from visual noise; (2) Existing edge-cloud partitioning methods overlook the step-wise redundancy unique to embodied tasks, thereby disrupting the physical continuity of motion. To address these issues, we propose a novel ECC inference framework, termed RAPID. Specifically, we developed an implementation tailored to the proposed framework. Experiments demonstrate this achieves a speedup of up to 1.73x with only 5%~7% overhead.

Keywords

Cite

@article{arxiv.2603.07949,
  title  = {RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models},
  author = {Zihao Zheng and Sicheng Tian and Hangyu Cao and Chenyue Li and Jiayu Chen and Maoliang Li and Xinhao Sun and Hailong Zou and Guojie Luo and Xiang Chen},
  journal= {arXiv preprint arXiv:2603.07949},
  year   = {2026}
}
R2 v1 2026-07-01T11:09:38.712Z