English

Habilis-$\beta$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model

Robotics 2026-02-24 v1 Machine Learning

Abstract

We introduce Habilis-β\beta, a fast-motion and long-lasting on-device vision-language-action (VLA) model designed for real-world deployment. Current VLA evaluation remains largely confined to single-trial success rates under curated resets, which fails to capture the fast-motion and long-lasting capabilities essential for practical operation. To address this, we introduce the Productivity-Reliability Plane (PRP), which evaluates performance through Tasks per Hour (TPH) and Mean Time Between Intervention (MTBI) under a continuous-run protocol that demands both high-speed execution and sustained robustness. Habilis-β\beta achieves high performance by integrating language-free pre-training on large-scale play data for robust interaction priors with post-training on cyclic task demonstrations that capture state drift across consecutive task iterations. The system further employs ESPADA for phase-adaptive motion shaping to accelerate free-space transit, utilizes rectified-flow distillation to enable high-frequency control on edge devices, and incorporates classifier-free guidance (CFG) as a deployment-time knob to dynamically balance instruction adherence and learned interaction priors. In 1-hour continuous-run evaluations, Habilis-β\beta achieves strong performance under the PRP metrics, compared to π0.5\pi_{0.5} in both simulation and real-world environments. In simulation, Habilis-β\beta achieves 572.6 TPH and 39.2 s MTBI (vs. 120.5 TPH and 30.5 s for π0.5\pi_{0.5}), while in a real-world humanoid logistics workflow it achieves 124 TPH and 137.4 s MTBI (vs. 19 TPH and 46.1 s for π0.5\pi_{0.5}). Finally, Habilis-β\beta achieves the highest reported performance on the standard RoboTwin 2.0 leaderboard across representative tasks, validating its effectiveness in complex manipulation scenarios.

Keywords

Cite

@article{arxiv.2602.18813,
  title  = {Habilis-$\beta$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model},
  author = {Tommoro Robotics and : and Jesoon Kang and Taegeon Park and Jisu An and Soo Min Kimm and Jaejoon Kim and Jinu Pahk and Byungju Kim and Junseok Lee and Namheon Baek and Sungwan Ha and Hojun Baek and Eduardo Ayerve Cruz and Wontae Kim and Junghyeon Choi and Yousuk Lee and Joonmo Han and Sunghyun Cho and Sunghyun Kwon and Soyoung Lee and Jun Ki Lee and Seung-Joon Yi and Byoung-Tak Zhang and Theo Taeyeong Kim},
  journal= {arXiv preprint arXiv:2602.18813},
  year   = {2026}
}