English
Related papers

Related papers: JoyAI-RA 0.1: A Foundation Model for Robotic Auton…

200 papers

While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning is limited by data scarcity and model capacity. To address…

Robotics · Computer Science 2025-10-15 Yi Yang , Kefan Gu , Yuqing Wen , Hebei Li , Yucheng Zhao , Tiancai Wang , Xudong Liu

Vision-Language-Action (VLA) models have shown remarkable achievements, driven by the rich implicit knowledge of their vision-language components. However, achieving generalist robotic agents demands precise grounding into physical…

Robotics · Computer Science 2025-07-15 Jialei Huang , Shuo Wang , Fanqi Lin , Yihang Hu , Chuan Wen , Yang Gao

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for open-world robot manipulation, but their practical deployment is often constrained by cost: billion-scale VLM backbones and iterative diffusion/flow-based action…

Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the…

Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is useful as a VLA initialization. In this paper, we study VLA…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Weifeng Lin , Siyuan Huang , Hao Li , Tingwei Chen , Ruichuan An , Xinyu Wei , Jianbo Liu , Hongsheng Li

Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environments. Although Vision-Language-Action (VLA) models provide strong perception and intuitive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Dongjie Huo , Haoyun Liu , Guoqing Liu , Dekang Qi , Zhiming Sun , Maoguo Gao , Jianxin He , Yandan Yang , Xinyuan Chang , Feng Xiong , Xing Wei , Zhiheng Ma , Mu Xu

Vision-Language-Action (VLA) models have emerged as promising solutions for robotic manipulation, yet their robustness to real-world physical variations remains critically underexplored. To bridge this gap, we propose Eva-VLA, the first…

Robotics · Computer Science 2026-03-17 Hanqing Liu , Shouwei Ruan , Jiahuan Long , Junqi Wu , Jiacheng Hou , Huili Tang , Tingsong Jiang , Weien Zhou , Wen Yao

Exploring open-world situations in an end-to-end manner is a promising yet challenging task due to the need for strong generalization capabilities. In particular, end-to-end autonomous driving in unstructured outdoor environments often…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hyunki Seong , Seongwoo Moon , Hojin Ahn , Jehun Kang , David Hyunchul Shim

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representations to connect…

Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires both high-level task planning (i.e., decomposing goals into…

Robotics · Computer Science 2025-06-03 Yi Yang , Jiaxuan Sun , Siqi Kou , Yihan Wang , Zhijie Deng

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread success of their precursors -- LLMs and VLMs. In this paper, we…

Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under…

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haitao Lin , Hanyang Yu , Jingshun Huang , He Zhang , Yonggen Ling , Ping Tan , Xiangyang Xue , Yanwei Fu

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Hao Luo , Yicheng Feng , Wanpeng Zhang , Sipeng Zheng , Ye Wang , Haoqi Yuan , Jiazheng Liu , Chaoyi Xu , Qin Jin , Zongqing Lu

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and…

The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained…

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene…

Robotics · Computer Science 2024-12-18 Qi Sun , Pengfei Hong , Tej Deep Pala , Vernon Toh , U-Xuan Tan , Deepanway Ghosal , Soujanya Poria

Mobile manipulation is the fundamental challenge for robotics to assist humans with diverse tasks and environments in everyday life. However, conventional mobile manipulation approaches often struggle to generalize across different tasks…

Robotics · Computer Science 2025-03-18 Zhenyu Wu , Yuheng Zhou , Xiuwei Xu , Ziwei Wang , Haibin Yan

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers…