English

EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation

Robotics 2025-11-10 v1 Computer Vision and Pattern Recognition

Abstract

While Vision-Language-Action (VLA) models map visual inputs and language instructions directly to robot actions, they often rely on costly hardware and struggle in novel or cluttered scenes. We introduce EverydayVLA, a 6-DOF manipulator that can be assembled for under $300, capable of modest payloads and workspace. A single unified model jointly outputs discrete and continuous actions, and our adaptive-horizon ensemble monitors motion uncertainty to trigger on-the-fly re-planning for safe, reliable operation. On LIBERO, EverydayVLA matches state-of-the-art success rates, and in real-world tests it outperforms prior methods by 49% in-distribution and 34.9% out-of-distribution. By combining a state-of-the-art VLA with cost-effective hardware, EverydayVLA democratizes access to a robotic foundation model and paves the way for economical use in homes and research labs alike. Experiment videos and details: https://everydayvla.github.io/

Keywords

Cite

@article{arxiv.2511.05397,
  title  = {EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation},
  author = {Samarth Chopra and Alex McMoil and Ben Carnovale and Evan Sokolson and Rajkumar Kubendran and Samuel Dickerson},
  journal= {arXiv preprint arXiv:2511.05397},
  year   = {2025}
}

Comments

Submitted to ICRA 2026

R2 v1 2026-07-01T07:26:27.244Z