中文
相关论文

相关论文: UrbanVLA: A Vision-Language-Action Model for Urban…

200 篇论文

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landscape remains highly fragmented and complex: as existing approaches vary substantially in…

机器人学 · 计算机科学 2026-04-14 Jinhui Ye , Ning Gao , Senqiao Yang , Jinliang Zheng , Zixuan Wang , Yuxin Chen , Pengguang Chen , Yilun Chen , Shu Liu , Jiaya Jia

Scaling Vision-Language-Action (VLA) models on large-scale data offers a promising path to achieving a more generalized driving intelligence. However, VLA models are limited by a ``supervision deficit'': the vast model capacity is…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Yingyan Li , Shuyao Shang , Weisong Liu , Bing Zhan , Haochen Wang , Yuqi Wang , Yuntao Chen , Xiaoman Wang , Yasong An , Chufeng Tang , Lu Hou , Lue Fan , Zhaoxiang Zhang

A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While large vision-language models (VLMs) trained on web-scale data…

In the domain of humanoid robot control, the fusion of Vision-Language-Action (VLA) with whole-body control is essential for semantically guided execution of real-world tasks. However, existing methods encounter challenges in terms of low…

机器人学 · 计算机科学 2026-03-06 Weikai Qin , Sichen Wu , Ci Chen , Mengfan Liu , Linxi Feng , Xinru Cui , Haoqi Han , Hesheng Wang

Vision Language Models (VLMs) bridge visual perception and linguistic reasoning. In Autonomous Driving (AD), this synergy has enabled Vision Language Action (VLA) models, which translate high-level multimodal understanding into driving…

机器人学 · 计算机科学 2026-03-11 Yuan Gao , Dengyuan Hua , Mattia Piccinini , Finn Rasmus Schäfer , Korbinian Moller , Lin Li , Johannes Betz

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To…

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation…

人工智能 · 计算机科学 2026-02-04 Qixin Zeng , Shuo Zhang , Hongyin Zhang , Renjie Wang , Han Zhao , Libang Zhao , Runze Li , Donglin Wang , Chao Huang

Urban research involves a wide range of scenarios and tasks that require the understanding of multi-modal data. Current methods often focus on specific data types and lack a unified framework in urban field for processing them…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Jie Feng , Shengyuan Wang , Tianhui Liu , Yanxin Xi , Yong Li

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is that the models exhibit…

机器人学 · 计算机科学 2026-02-05 Guoqing Ma , Siheng Wang , Zeyu Zhang , Shan Yu , Hao Tang

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings without explicitly…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Weiye Zhu , Zekai Zhang , Xiangchen Wang , Hewei Pan , Teng Wang , Tiantian Geng , Rongtao Xu , Feng Zheng

Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics,…

机器人学 · 计算机科学 2026-05-18 Esteban Padilla-Cerdio , Boyang Sun , Marc Pollefeys , Hermann Blum

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor…

机器人学 · 计算机科学 2025-12-23 Xu Liu , Yu Liu , Hanshuo Qiu , Yang Qirong , Zhouhui Lian

Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructions into robotic actions. Despite their recent advancements,…

机器人学 · 计算机科学 2025-05-27 Tuan Van Vo , Tan Quang Nguyen , Khang Minh Nguyen , Duy Ho Minh Nguyen , Minh Nhat Vu

We propose LCLA (Language-Conditioned Latent Alignment), a framework for vision-language navigation that learns modular perception-action interfaces by aligning sensory observations to a latent representation of an expert policy. The expert…

机器人学 · 计算机科学 2026-02-11 Nitesh Subedi , Adam Haroon , Samuel Tetteh , Prajwal Koirala , Cody Fleming , Soumik Sarkar

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov…

机器学习 · 计算机科学 2026-04-13 Lei Xiao , Jifeng Li , Juntao Gao , Feiyang Ye , Yan Jin , Jingjing Qian , Jing Zhang , Yong Wu , Xiaoyuan Yu

Vision-Language-Action (VLA) models offer promising capabilities for autonomous driving through multimodal understanding. However, their utilization in safety-critical scenarios is constrained by inherent limitations, including imprecise…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Yiru Wang , Zichong Gu , Yu Gao , Anqing Jiang , Zhigang Sun , Shuo Wang , Yuwen Heng , Hao Sun

General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models…

机器人学 · 计算机科学 2026-02-26 Yue Yang , Shuo Cheng , Yu Fang , Homanga Bharadhwaj , Mingyu Ding , Gedas Bertasius , Daniel Szafir

Vision-language-action models have emerged as a crucial paradigm in robotic manipulation. However, existing VLA models exhibit notable limitations in handling ambiguous language instructions and unknown environmental states. Furthermore,…

机器人学 · 计算机科学 2025-08-26 Helong Huang , Min Cen , Kai Tan , Xingyue Quan , Guowei Huang , Hong Zhang

Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Zhaoshu Yu , Bo Wang , Pengpeng Zeng , Haonan Zhang , Ji Zhang , Zheng Wang , Lianli Gao , Jingkuan Song , Nicu Sebe , Heng Tao Shen