English

Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning

Computer Vision and Pattern Recognition 2026-02-25 v2 Artificial Intelligence Machine Learning Robotics

Abstract

Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy reasoning traces. We propose Fast-ThinkAct, an efficient reasoning framework that achieves compact yet performant planning through verbalizable latent reasoning. Fast-ThinkAct learns to reason efficiently with latent CoTs by distilling from a teacher, driven by a preference-guided objective to align manipulation trajectories that transfers both linguistic and visual planning capabilities for embodied control. This enables reasoning-enhanced policy learning that effectively connects compact reasoning to action execution. Extensive experiments across diverse embodied manipulation and reasoning benchmarks demonstrate that Fast-ThinkAct achieves strong performance with up to 89.3% reduced inference latency over state-of-the-art reasoning VLAs, while maintaining effective long-horizon planning, few-shot adaptation, and failure recovery.

Keywords

Cite

@article{arxiv.2601.09708,
  title  = {Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning},
  author = {Chi-Pin Huang and Yunze Man and Zhiding Yu and Min-Hung Chen and Jan Kautz and Yu-Chiang Frank Wang and Fu-En Yang},
  journal= {arXiv preprint arXiv:2601.09708},
  year   = {2026}
}

Comments

CVPR 2026. Project page: https://jasper0314-huang.github.io/fast-thinkact/

R2 v1 2026-07-01T09:04:42.591Z