English
Related papers

Related papers: How VLAs (Really) Work In Open-World Environments

200 papers

Vision-language action (VLA) policies often report strong manipulation benchmark performance with relatively few demonstrations, but it remains unclear whether this reflects robust language-to-object grounding or reliance on…

Robotics · Computer Science 2026-03-02 David Emukpere , Romain Deffayet , Jean-Michel Renders

Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow language. When presented with instructions that lack strong scene-specific supervision, VLAs…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Yu Fang , Yuchun Feng , Dong Jing , Jiaqi Liu , Yue Yang , Zhenyu Wei , Daniel Szafir , Mingyu Ding

Recent vision-language-action models (VLAs) build upon pretrained vision-language models and leverage diverse robot datasets to demonstrate strong task execution, language following ability, and semantic generalization. Despite these…

Robotics · Computer Science 2025-04-29 Moo Jin Kim , Chelsea Finn , Percy Liang

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in…

Achieving truly adaptive embodied intelligence requires agents that learn not just by imitating static demonstrations, but by continuously improving through environmental interaction, which is akin to how humans master skills through…

Robotics · Computer Science 2025-12-17 Zechen Bai , Chen Gao , Mike Zheng Shou

Large-scale generative models are shown to be useful for sampling meaningful candidate solutions, yet they often overlook task constraints and user preferences. Their full power is better harnessed when the models are coupled with external…

Artificial Intelligence · Computer Science 2024-08-13 Lin Guan , Yifan Zhou , Denis Liu , Yantian Zha , Heni Ben Amor , Subbarao Kambhampati

Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation…

Artificial Intelligence · Computer Science 2026-02-04 Qixin Zeng , Shuo Zhang , Hongyin Zhang , Renjie Wang , Han Zhao , Libang Zhao , Runze Li , Donglin Wang , Chao Huang

In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broader multi-modal perturbations that arise in actions,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Jianing Guo , Zhenhong Wu , Chang Tu , Yiyao Ma , Xiangqi Kong , Zhiqian Liu , Jiaming Ji , Shuning Zhang , Yuanpei Chen , Kai Chen , Qi Dou , Yaodong Yang , Xianglong Liu , Huijie Zhao , Weifeng Lv , Simin Li

Current vision-language-action (VLA) models, pre-trained on large-scale robotic data, exhibit strong multi-task capabilities and generalize well to variations in visual and language instructions for manipulation. However, their success rate…

Robotics · Computer Science 2025-10-17 Han Zhao , Jiaxuan Zhang , Wenxuan Song , Pengxiang Ding , Donglin Wang

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates…

Vision-Language Models (VLMs) are powerful yet computationally intensive for widespread practical deployments. To address such challenge without costly re-training, post-training acceleration techniques like quantization and token reduction…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yizheng Sun , Hao Li , Chang Xu , Hongpeng Zhou , Chenghua Lin , Riza Batista-Navarro , Jingyuan Sun

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for open-world robot manipulation, but their practical deployment is often constrained by cost: billion-scale VLM backbones and iterative diffusion/flow-based action…

Leveraging temporal context is crucial for success in partially observable robotic tasks. However, prior work in behavior cloning has demonstrated inconsistent performance gains when using multi-frame observations. In this paper, we…

Robotics · Computer Science 2025-10-07 Huiwon Jang , Sihyun Yu , Heeseung Kwon , Hojin Jeon , Younggyo Seo , Jinwoo Shin

Vision-Language-Action (VLA) models trained on large robot datasets promise general-purpose, robust control across diverse domains and embodiments. However, existing approaches often fail out-of-the-box when deployed in novel environments,…

Robotics · Computer Science 2025-10-21 Ruihan Zhao , Tyler Ingebrand , Sandeep Chinchali , Ufuk Topcu

Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as LLaVa, InstructBLIP, and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Siddharth Karamcheti , Suraj Nair , Ashwin Balakrishna , Percy Liang , Thomas Kollar , Dorsa Sadigh

Vision-Language-Action (VLA) models improve action generation by conditioning policies on rich vision-language information. However, current auto-regressive policies are constrained by three bottlenecks: (1) architectural bias drives models…

Robotics · Computer Science 2026-03-31 Yichi Zhang , Weihao Yuan , Yizhuo Zhang , Xidong Zhang , Jia Wan

Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental…

Robotics · Computer Science 2025-02-10 Hong Lu , Hengxu Li , Prithviraj Singh Shahani , Stephanie Herbers , Matthias Scheutz

While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deployment remains challenged by uncertainty in learning-based action generation. Low-quality…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Zhen Sun , Yongjian Guo , Haoran Sun , Luqiao Wang , Wei Lu , Jiachi Ji , Shengzhe Ji , Junwu Xiong , Zhijun Meng

Vision-Language-Action (VLA) models have emerged as powerful generalist policies for robotic control, yet their performance scaling across model architectures and hardware platforms, as well as their associated power budgets, remain poorly…

Artificial Intelligence · Computer Science 2026-01-27 Amir Taherin , Juyi Lin , Arash Akbari , Arman Akbari , Pu Zhao , Weiwei Chen , David Kaeli , Yanzhi Wang

Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between…

Robotics · Computer Science 2026-04-21 Haiweng Xu , Sipeng Zheng , Hao Luo , Wanpeng Zhang , Ziheng Xi , Zongqing Lu
‹ Prev 1 3 4 5 6 7 10 Next ›