English
Related papers

Related papers: ReFineVLA: Reasoning-Aware Teacher-Guided Transfer…

200 papers

Vision-language-action (VLA) models that directly predict multi-step action chunks from current observations face inherent limitations due to constrained scene understanding and weak future anticipation capabilities. In contrast, video…

Vision-Language-Action (VLA) models have shown remarkable success in robotic tasks like manipulation by fusing a language model's reasoning with a vision model's 3D understanding. However, their high computational cost remains a major…

Robotics · Computer Science 2026-03-18 Zebin Yang , Yijiahao Qi , Tong Xie , Bo Yu , Shaoshan Liu , Meng Li

Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial…

Robotics · Computer Science 2025-10-22 Siyu Xu , Yunke Wang , Chenghao Xia , Dihao Zhu , Tao Huang , Chang Xu

Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact…

Robotics · Computer Science 2026-02-16 Yike Zhang , Yaonan Wang , Xinxin Sun , Kaizhen Huang , Zhiyuan Xu , Junjie Ji , Zhengping Che , Jian Tang , Jingtao Sun

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores. On the other…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Yifan Li , Yikai Wang , Yanwei Fu , Dongyu Ru , Zheng Zhang , Tong He

Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires further fine-tuning -- for example, via reinforcement…

Machine Learning · Computer Science 2026-05-04 Charles Xu , Jost Tobias Springenberg , Michael Equi , Ali Amin , Adnan Esmail , Sergey Levine , Liyiming Ke

The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into active agents for…

Robotics · Computer Science 2025-11-11 Dapeng Zhang , Jing Sun , Chenghui Hu , Xiaoyan Wu , Zhenlong Yuan , Rui Zhou , Fei Shen , Qingguo Zhou

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for generalist robotic control. Built upon vision-language model (VLM) architectures, VLAs predict actions conditioned on visual observations and language…

Robotics · Computer Science 2026-05-26 Weikang Qiu , Huashuo Lei , Tinglin Huang , Rex Ying

Reinforcement Learning (RL) has shown promise in improving the reasoning abilities of Large Language Models (LLMs). However, the specific challenges of adapting RL to multimodal data and formats remain relatively unexplored. In this work,…

Machine Learning · Computer Science 2025-05-20 Zirun Guo , Minjie Hong , Tao Jin

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yuhao Zhou , Yunpeng Zhu , Yang Zhou , Jindi Lyu , Jian Lan , Zhangyuan Wang , Dan Si , Thomas Seidl , Qing Ye , Jiancheng Lyu

Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their large-scale training, VLAs are often brittle to…

Robotics · Computer Science 2024-10-04 Asher J. Hancock , Allen Z. Ren , Anirudha Majumdar

Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA,…

Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a critical obstacle to real-world deployment. Improving…

Robotics · Computer Science 2026-02-03 Yujie Wei , Jiahan Fan , Jiyu Guo , Ruichen Zhen , Rui Shao , Xiu Su , Zeke Xie , Shuo Yang

Vision language action (VLA) models enable generalist robotic agents but often exhibit language ignorance, relying on visual shortcuts and remaining insensitive to instruction changes. We present Prospective Grounding and Alignment VLA…

Robotics · Computer Science 2026-04-14 Nastaran Darabi , Amit Ranjan Trivedi

Vision-language-action models (VLAs) trained on large-scale robotic datasets have demonstrated strong performance on manipulation tasks, including bimanual tasks. However, because most public datasets focus on single-arm demonstrations,…

Robotics · Computer Science 2026-02-24 Hokyun Im , Euijin Jeong , Andrey Kolobov , Jianlong Fu , Youngwoon Lee

The emergence of vision-language-action (VLA) models has given rise to foundation models for robot manipulation. Although these models have achieved significant improvements, their generalization in multi-task manipulation remains limited.…

We introduce iFlyBot-VLA, a large-scale Vision-Language-Action (VLA) model trained under a novel framework. The main contributions are listed as follows: (1) a latent action model thoroughly trained on large-scale human and robotic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Yuan Zhang , Chenyu Xue , Wenjie Xu , Chao Ji , Jiajia wu , Jia Pan

Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental…

Robotics · Computer Science 2025-02-10 Hong Lu , Hengxu Li , Prithviraj Singh Shahani , Stephanie Herbers , Matthias Scheutz

Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and…

Robotics · Computer Science 2026-05-29 Shengyu Si , Yuanzhuo Lu , Ruimeng Yang , Ziyi Ye , Zuxuan Wu , Yu-Gang Jiang

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yicheng Liu , Shiduo Zhang , Zibin Dong , Baijun Ye , Tianyuan Yuan , Xiaopeng Yu , Linqi Yin , Chenhao Lu , Junhao Shi , Luca Jiang-Tao Yu , Liangtao Zheng , Tao Jiang , Jingjing Gong , Xipeng Qiu , Hang Zhao