English
Related papers

Related papers: FutureVLA: Joint Visuomotor Prediction for Vision-…

200 papers

A generalist robot should perform effectively across various environments. However, most existing approaches heavily rely on scaling action-annotated data to enhance their capabilities. Consequently, they are often limited to single…

Robotics · Computer Science 2025-11-04 Qingwen Bu , Yanting Yang , Jisong Cai , Shenyuan Gao , Guanghui Ren , Maoqing Yao , Ping Luo , Hongyang Li

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates…

Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact, occlusion, and object motion, and the geometry that later…

Robotics · Computer Science 2026-05-22 Chushan Zhang , Ruihan Lu , Jinguang Tong , Xuesong Li , Yikai Wang , Hongdong Li

Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak…

Robotics · Computer Science 2026-05-29 Zhongyu Xia , Yousen Tang , Bingqing Wei , Yongtao Wang

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models…

Robotics · Computer Science 2026-04-06 Jiashu Yang , Yifan Han , Yucheng Xie , Ning Guo , Wenzhao Lian

Predicting future frames of a video sequence has been a problem of high interest in the field of Computer Vision as it caters to a multitude of applications. The ability to predict, anticipate and reason about future events is the essence…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Jasmeen Kaur , Sukhendu Das

Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructions into desired robotic actions. Despite their…

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xinyang Wang , Qian Liu , Wenjie Ding , Zhao Yang , Wei Li , Chang Liu , Bailin Li , Kun Zhan , Xianpeng Lang , Wei Chen

Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data,…

Robotics · Computer Science 2025-08-05 Cui Miao , Tao Chang , Meihan Wu , Hongbin Xu , Chun Li , Ming Li , Xiaodong Wang

Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual representations of most VLA models are often dominated by global object appearance and struggle…

Robotics · Computer Science 2026-05-19 Weijie Kong , Zhian Su , Wei Yu , Huixu Dong

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yuhao Zhou , Yunpeng Zhu , Yang Zhou , Jindi Lyu , Jian Lan , Zhangyuan Wang , Dan Si , Thomas Seidl , Qing Ye , Jiancheng Lyu

Vision-Language-Action (VLA) models built on pretrained Vision-Language Models (VLMs) show strong potential but are limited in practicality due to their large parameter counts. To mitigate this issue, using a lightweight VLM has been…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Chaojun Ni , Cheng Chen , Xiaofeng Wang , Zheng Zhu , Wenzhao Zheng , Boyuan Wang , Tianrun Chen , Guosheng Zhao , Haoyun Li , Zhehao Dong , Qiang Zhang , Yun Ye , Yang Wang , Guan Huang , Wenjun Mei

In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, trajectory, orientation, and relative displacement) from…

Robotics · Computer Science 2026-03-19 Gaoge Han , Zhengqing Gao , Ziwen Li , Jiaxin Huang , Shaoli Huang , Fakhri Karray , Mingming Gong , Tongliang Liu

Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, their real-world deployment remains slow and inefficient: demonstration videos are often sped up by 5-10x to appear smooth, with…

Robotics · Computer Science 2025-12-02 Jiaming Tang , Yufei Sun , Yilong Zhao , Shang Yang , Yujun Lin , Zhuoyang Zhang , James Hou , Yao Lu , Zhijian Liu , Song Han

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and…

Robotics · Computer Science 2025-10-28 Anqi Li , Zhiyong Wang , Jiazhao Zhang , Minghan Li , Yunpeng Qi , Zhibo Chen , Zhizheng Zhang , He Wang

Vision-language-action (VLA) models have significantly advanced robotic manipulation by integrating vision-language models (VLMs), and action decoders into a unified architecture. However, their deployment on resource-constrained edge…

Robotics · Computer Science 2025-10-30 Jiahong Chen , Jing Wang , Long Chen , Chuwei Cai , Jinghui Lu

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions,…

Robotics · Computer Science 2026-02-17 Jingwen Sun , Wenyao Zhang , Zekun Qi , Shaojie Ren , Zezhi Liu , Hanxin Zhu , Guangzhong Sun , Xin Jin , Zhibo Chen

Vision-language-action (VLA) models provide a powerful approach to training control policies for physical systems, such as robots, by combining end-to-end learning with transfer of semantic knowledge from web-scale vision-language model…

A reliable driving assistant should provide consistent responses based on temporally grounded reasoning derived from observed information. In this work, we investigate whether Vision-Language Models (VLMs), when applied as driving…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Chun-Peng Chang , Chen-Yu Wang , Holger Caesar , Alain Pagani
‹ Prev 1 4 5 6 7 8 10 Next ›