English
Related papers

Related papers: Experiences from Benchmarking Vision-Language-Acti…

200 papers

General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models…

Robotics · Computer Science 2026-02-26 Yue Yang , Shuo Cheng , Yu Fang , Homanga Bharadhwaj , Mingyu Ding , Gedas Bertasius , Daniel Szafir

In this study, we address the problem of language-guided robotic manipulation, where a robot is required to manipulate a wide range of objects based on visual observations and natural language instructions. This task is essential for…

Robotics · Computer Science 2026-03-17 Yusuke Takagi , Motonari Kambara , Daichi Yashima , Koki Seno , Kento Tokura , Komei Sugiura

Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protocols aligned with human preferences. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuhan Hao , Zhengning Li , Lei Sun , Weilong Wang , Naixin Yi , Sheng Song , Caihong Qin , Mofan Zhou , Yifei Zhan , Xianpeng Lang

Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising generalization in manipulation tasks. However, most existing approaches primarily rely on direct…

Robotics · Computer Science 2026-03-02 Jiasong Xiao , Yutao She , Kai Li , Yuyang Sha , Ziang Cheng , Ziang Tong

Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representations to connect…

The evaluation of Vision-Language-Action (VLA) agents is hindered by the coarse, end-task success metric that fails to provide precise skill diagnosis or measure robustness to real-world perturbations. This challenge is exacerbated by a…

Robotics · Computer Science 2025-10-22 Jierui Peng , Yanyan Zhang , Yicheng Duan , Tuo Liang , Vipin Chaudhary , Yu Yin

Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under…

Machine Learning · Computer Science 2026-01-15 Jijia Liu , Feng Gao , Bingwen Wei , Xinlei Chen , Qingmin Liao , Yi Wu , Chao Yu , Yu Wang

Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimization. This "flat" training paradigm, inherited from language…

Robotics · Computer Science 2026-05-29 Daojie Peng , Fulong Ma , Jiahang Cao , Qiang Zhang , Xupeng Xie , Jian Guo , Ping Luo , Andrew F. Luo , Boyu Zhou , Jun Ma

Robotic manipulation is a fundamental component of automation. However, traditional perception-planning pipelines often fall short in open-ended tasks due to limited flexibility, while the architecture of a single end-to-end…

Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved…

A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation tasks typically assume that instructions are perfectly aligned…

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is the lack of robotic data,…

Confidence estimation for Vision-Language-Action (VLA) models is essential for robots to perform manipulation tasks in the open world, providing crucial signals for risk-sensitive decision-making and failure anticipation. Existing…

Robotics · Computer Science 2026-05-29 Dehao Huang , Aoxiang Gu , Chengjie Zhang , Bolin Zou , Wenlong Dong , Zilang Cen , Yue Wang , Hong Zhang

Vision-Language-Action (VLA) models aim to provide a single generalist controller for robots, but today's systems fall short on the criteria that matter for real-world deployment. Frontier models are closed, open-weight alternatives are…

Building generalist embodied agents requires integrating perception, language understanding, and action, which are core capabilities addressed by Vision-Language-Action (VLA) approaches based on multimodal foundation models, including…

Robotics · Computer Science 2026-04-08 StarVLA Community

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Jianke Zhang , Yanjiang Guo , Yucheng Hu , Xiaoyu Chen , Xiang Zhu , Jianyu Chen

Vision-Language Models (VLMs) are increasingly pivotal for generalist robot manipulation, enabling tasks such as physical reasoning, policy generation, and failure detection. However, their proficiency in these high-level applications often…

Robotics · Computer Science 2025-07-01 Atharva Gundawar , Som Sagar , Ransalu Senanayake

While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning is limited by data scarcity and model capacity. To address…

Robotics · Computer Science 2025-10-15 Yi Yang , Kefan Gu , Yuqing Wen , Hebei Li , Yucheng Zhao , Tiancai Wang , Xudong Liu

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Anupam Pani , Yanchao Yang

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control.…