English
Related papers

Related papers: See Once, Then Act: Vision-Language-Action Model w…

200 papers

Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene through multiple…

Robotics · Computer Science 2025-12-23 Jin Wang , Kim Tien Ly , Jacques Cloete , Nikos Tsagarakis , Ioannis Havoutis

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely…

Interactive robot learning is a challenging problem as the robot is present with human users who expect the robot to learn novel skills to solve novel tasks perpetually with sample efficiency. In this work we present a framework for robots…

Robotics · Computer Science 2026-03-31 Weiwei Gu , Suresh Kondepudi , Anmol Gupta , Lixiao Huang , Nakul Gopalan

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentalize perception, planning, and decision-making, simplifying…

Robotics · Computer Science 2025-02-05 Pengxiang Ding , Han Zhao , Wenjie Zhang , Wenxuan Song , Min Zhang , Siteng Huang , Ningxi Yang , Donglin Wang

Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising generalization in manipulation tasks. However, most existing approaches primarily rely on direct…

Robotics · Computer Science 2026-03-02 Jiasong Xiao , Yutao She , Kai Li , Yuyang Sha , Ziang Cheng , Ziang Tong

Vision-language action (VLA) policies often report strong manipulation benchmark performance with relatively few demonstrations, but it remains unclear whether this reflects robust language-to-object grounding or reliance on…

Robotics · Computer Science 2026-03-02 David Emukpere , Romain Deffayet , Jean-Michel Renders

Vision-Language Models (VLMs) have emerged as a promising approach to address the data scarcity challenge in robotics, enabling the development of generalizable visuomotor control policies. While models like OpenVLA showcase the potential…

Vision-Language-Action (VLA) models have recently made significant advance in multi-task, end-to-end robotic control, due to the strong generalization capabilities of Vision-Language Models (VLMs). A fundamental challenge in developing such…

Robotics · Computer Science 2025-06-17 Yuqing Wen , Kefan Gu , Haoxuan Liu , Yucheng Zhao , Tiancai Wang , Haoqiang Fan , Xiaoyan Sun

Pre-trained Vision-Language-Action (VLA) models have achieved remarkable success in improving robustness and generalization for end-to-end robotic manipulation. However, these models struggle with long-horizon tasks due to their lack of…

Robotics · Computer Science 2025-11-13 Runhao Li , Wenkai Guo , Zhenyu Wu , Changyuan Wang , Haoyuan Deng , Zhenyu Weng , Yap-Peng Tan , Ziwei Wang

This study evaluates two leading approaches for teaching construction robots new skills to understand their applicability for construction automation: a Vision-Language-Action (VLA) model and Reinforcement Learning (RL) methods. The goal is…

Robotics · Computer Science 2026-03-02 Zhaofeng Hu , Hongrui Yu , Vaidhyanathan Chandramouli , Ci-Jyun Liang

General-purpose robots capable of performing diverse tasks require synergistic reasoning and acting capabilities. However, recent dual-system approaches, which separate high-level reasoning from low-level acting, often suffer from…

Robotics · Computer Science 2026-03-03 Fanqi Lin , Ruiqian Nai , Yingdong Hu , Jiacheng You , Junming Zhao , Yang Gao

Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories without predictive…

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities…

Robotics · Computer Science 2025-06-02 Zhongyi Zhou , Yichen Zhu , Junjie Wen , Chaomin Shen , Yi Xu

Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environments. In recent…

Robotics · Computer Science 2025-09-04 Rui Shao , Wei Li , Lingsen Zhang , Renshan Zhang , Zhiyang Liu , Ran Chen , Liqiang Nie

Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. A key bottleneck is that collecting robotic data often requires expertise or specialized hardware, limiting accessibility and…

Robotics · Computer Science 2025-05-13 Gi-Cheon Kang , Junghyun Kim , Kyuhwan Shim , Jun Ki Lee , Byoung-Tak Zhang

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for robotics require finetuning to contend with varying…

Traditional robotic systems typically decompose intelligence into independent modules for computer vision, natural language processing, and motion control. Vision-Language-Action (VLA) models fundamentally transform this approach by…

Robotics · Computer Science 2025-08-12 Heran Wu , Zirun Zhou , Jingfeng Zhang

Vision-Language-Action (VLA) models have recently emerged as a powerful paradigm for robotic manipulation. Despite substantial progress enabled by large-scale pretraining and supervised fine-tuning (SFT), these models face two fundamental…