English
Related papers

Related papers: DM0: An Embodied-Native Vision-Language-Action Mod…

200 papers

The development of foundation models for embodied intelligence critically depends on access to large-scale, high-quality robot demonstration data. Recent approaches have sought to address this challenge by training on large collections of…

Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation~(VLN) which requires visual and natural…

Artificial Intelligence · Computer Science 2024-01-25 Raphael Schumann , Wanrong Zhu , Weixi Feng , Tsu-Jui Fu , Stefan Riezler , William Yang Wang

Vision-Language-Action (VLA) models trained on large robot datasets promise general-purpose, robust control across diverse domains and embodiments. However, existing approaches often fail out-of-the-box when deployed in novel environments,…

Robotics · Computer Science 2025-10-21 Ruihan Zhao , Tyler Ingebrand , Sandeep Chinchali , Ufuk Topcu

Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inherently challenging as it requires both accurate target…

Robotics · Computer Science 2025-05-30 Shaoan Wang , Jiazhao Zhang , Minghan Li , Jiahang Liu , Anqi Li , Kui Wu , Fangwei Zhong , Junzhi Yu , Zhizheng Zhang , He Wang

We introduce iFlyBot-VLA, a large-scale Vision-Language-Action (VLA) model trained under a novel framework. The main contributions are listed as follows: (1) a latent action model thoroughly trained on large-scale human and robotic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Yuan Zhang , Chenyu Xue , Wenjie Xu , Chao Ji , Jiajia wu , Jia Pan

Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environments. In recent…

Robotics · Computer Science 2025-09-04 Rui Shao , Wei Li , Lingsen Zhang , Renshan Zhang , Zhiyang Liu , Ran Chen , Liqiang Nie

Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training paradigms face a critical trade-off: joint end-to-end training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yuyuan Yang , Junkun Hong , Hongrong Wang , Honghao Cai , Xunpeng Ren , Ge Wang , Mingcong Lei , Shenhao Yan , Jiahao Yang , Chengsi Yao , Xi Li , Yiming Zhao , Yatong Han , Jinke Ren

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jiazhao Zhang , Kunyu Wang , Rongtao Xu , Gengze Zhou , Yicong Hong , Xiaomeng Fang , Qi Wu , Zhizheng Zhang , He Wang

While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound reasoning capabilities inherited from their base…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tianshuo Yang , Guanyu Chen , Yutian Chen , Zhixuan Liang , Yitian Liu , Zanxin Chen , Chunpu Xu , Haotian Liang , Jiangmiao Pang , Yao Mu , Ping Luo

Recent vision-language-action (VLA) systems have demonstrated strong capabilities in embodied manipulation. However, most existing VLA policies rely on limited observation windows and end-to-end action prediction, which makes them brittle…

Robotics · Computer Science 2026-04-16 Zhen Liu , Xinyu Ning , Zhe Hu , Xinxin Xie , Weize Li , Zhipeng Tang , Chongyu Wang , Zejun Yang , Hanlin Wang , Yitong Liu , Zhongzhu Pu

Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision-language-action…

Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular…

Vision-language-action (VLA) models are emerging as embodied foundation models for robotic manipulation, but their deployment introduces a new unlearning challenge: removing unsafe, spurious, or privacy-sensitive behaviors without degrading…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Ravi Ranjan , Agoritsa Polyzou

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates…

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall far short of…

Robotics · Computer Science 2025-09-03 Bear Häon , Kaylene Stocking , Ian Chuang , Claire Tomlin

The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked diffusion models, a paradigm distinct from autoregressive…

Robotics · Computer Science 2025-09-11 Yuqing Wen , Hebei Li , Kefan Gu , Yucheng Zhao , Tiancai Wang , Xiaoyan Sun

Vision-Language-Action models (VLAs) hold immense promise for enabling generalist robot manipulation. However, the best way to build them remains an open question. Current approaches often add complexity, such as modifying the existing…

Robotics · Computer Science 2025-10-16 Ankit Goyal , Hugo Hadfield , Xuning Yang , Valts Blukis , Fabio Ramos

Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of explicit mechanisms for multimodal reasoning and anticipating…

The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping…

Robotics · Computer Science 2026-04-29 Yi Chen , Yuying Ge , Hui Zhou , Mingyu Ding , Yixiao Ge , Xihui Liu

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodiment-specific…

Robotics · Computer Science 2026-05-26 Boyu Li , Chaoyi Xu , Haoqi Yuan , Xinrun Xu , Börje F. Karlsson , Dongbin Zhao , Haoran Li , Zongqing Lu