English
Related papers

Related papers: X-DiffVLA: X-Embodied Diffusion Action Heads for V…

200 papers

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control.…

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation,…

Robotics · Computer Science 2026-04-07 Yifu Yuan , Haiqin Cui , Yaoting Huang , Yibin Chen , Fei Ni , Zibin Dong , Pengyi Li , Yan Zheng , Hongyao Tang , Jianye Hao

The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentalize perception, planning, and decision-making, simplifying…

Robotics · Computer Science 2025-02-05 Pengxiang Ding , Han Zhao , Wenjie Zhang , Wenxuan Song , Min Zhang , Siteng Huang , Ningxi Yang , Donglin Wang

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haitao Lin , Hanyang Yu , Jingshun Huang , He Zhang , Yonggen Ling , Ping Tan , Xiangyang Xue , Yanwei Fu

Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this survey, we argue that future advances in VLA will depend…

Robotic Vision-Language-Action (VLA) models generalize well for open-ended manipulation, but their perception is fragile under sensing-stage degradations such as extreme low light, motion blur, and black clipping. We present E-VLA, an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Jiajun Zhai , Hao Shi , Shangwei Guo , Kailun Yang , Kaiwei Wang

Vision-Language-Action (VLA) models have demonstrated significant advantages in robotic manipulation. However, their reliance on vision and language often leads to suboptimal performance in tasks involving visual occlusion, fine-grained…

Cross-embodiment manipulation is crucial for enhancing the scalability of robot manipulation and reducing the high cost of data collection. However, the significant differences between embodiments, such as variations in action spaces and…

Robotics · Computer Science 2026-03-17 Juncheng Mu , Sizhe Yang , Hojin Bae , Feiyu Jia , Qingwei Ben , Boyi Li , Huazhe Xu , Jiangmiao Pang

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements.…

Vision-Language-Action (VLA) models are promising for generalist robot manipulation but remain brittle in out-of-distribution (OOD) settings, especially with limited real-robot data. To resolve the generalization bottleneck, we introduce a…

Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and continuous-action flow-matching (e.g. pi-0.5). Yet preference alignment via Direct…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Zhi Liu

Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged as promising candidates for end-to-end autonomous driving. However, these models typically face challenges in inference latency, action precision, and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Jiaru Zhang , Manav Gagvani , Can Cui , Juntong Peng , Ruqi Zhang , Ziran Wang

Achieving human-like dexterous manipulation remains a major challenge for general-purpose robots. While Vision-Language-Action (VLA) models show potential in learning skills from demonstrations, their scalability is limited by scarce…

Robotics · Computer Science 2025-12-16 Yu Cui , Yujian Zhang , Lina Tao , Yang Li , Xinyu Yi , Zhibin Li

Robots deployed in dynamic environments must be able to not only follow diverse language instructions but flexibly adapt when user intent changes mid-execution. While recent Vision-Language-Action (VLA) models have advanced multi-task…

Robotics · Computer Science 2025-06-05 Meng Li , Zhen Zhao , Zhengping Che , Fei Liao , Kun Wu , Zhiyuan Xu , Pei Ren , Zhao Jin , Ning Liu , Jian Tang

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches…

Robotics · Computer Science 2025-09-19 Anzhe Chen , Yifei Yang , Zhenjie Zhu , Kechun Xu , Zhongxiang Zhou , Rong Xiong , Yue Wang

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose embodied intelligence. However, current SOTA VLAs are…

Robotics · Computer Science 2025-10-10 Yandu Chen , Kefan Gu , Yuqing Wen , Yucheng Zhao , Tiancai Wang , Liqiang Nie

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained…

Robotics · Computer Science 2025-12-09 Yichao Shen , Fangyun Wei , Zhiying Du , Yaobo Liang , Yan Lu , Jiaolong Yang , Nanning Zheng , Baining Guo

Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical…

Acting in human environments is a crucial capability for general-purpose robots, necessitating a robust understanding of natural language and its application to physical tasks. This paper seeks to harness the capabilities of diffusion…

Robotics · Computer Science 2026-04-28 Jonas Bode , Raphael Memmesheimer , Sven Behnke

Building generalist embodied agents requires integrating perception, language understanding, and action, which are core capabilities addressed by Vision-Language-Action (VLA) approaches based on multimodal foundation models, including…

Robotics · Computer Science 2026-04-08 StarVLA Community
‹ Prev 1 8 9 10 Next ›