English
Related papers

Related papers: Listen, Look, Drive: Coupling Audio Instructions f…

200 papers

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates…

Machine Learning · Computer Science 2019-04-09 Khanh Nguyen , Debadeepta Dey , Chris Brockett , Bill Dolan

Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique object-level multimodal LLM architecture that merges vectorized numeric modalities…

Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face challenges in contact-rich manipulation scenarios that require precise physical…

Robotics · Computer Science 2026-05-19 Xiaoqi Li , Muhe Cai , Jiadong Xu , Juan Zhu , Hongwei Fan , Yan Shen , Guangrui Ren , Hao Dong

Long-horizon robotic manipulation remains challenging for Vision-Language-Action (VLA) models despite recent progress in zero-shot generalization and simulation-to-real-world transfer. Current VLA models suffer from stage hallucination,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Zeting Liu , Zida Yang , Zeyu Zhang , Hao Tang

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often leading to…

Robotics · Computer Science 2025-09-10 Qi Lv , Weijie Kong , Hao Li , Jia Zeng , Zherui Qiu , Delin Qu , Haoming Song , Qizhi Chen , Xiang Deng , Jiangmiao Pang

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Binjie Zhang , Mike Zheng Shou

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We…

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method that can effectively…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Chenxu Dang , Jie Wang , Guang Li , Zhiwen Hou , Zihan You , Hangjun Ye , Jie Ma , Long Chen , Yan Wang

Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-tuning of pre-trained Vision-Language Models (VLMs), feeding…

Robotics · Computer Science 2026-02-04 Kun Wang , Xiao Feng , Mingcheng Qu , Tonghua Su

End-to-end autonomous driving frameworks face persistent challenges in generalization, training efficiency, and interpretability. While recent methods leverage Vision-Language Models (VLMs) through supervised learning on large-scale…

Robotics · Computer Science 2025-12-11 Lin Li , Yuxin Cai , Jianwu Fang , Jianru Xue , Chen Lv

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Tianwen Qian , Jingjing Chen , Linhai Zhuo , Yang Jiao , Yu-Gang Jiang

Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their large-scale training, VLAs are often brittle to…

Robotics · Computer Science 2024-10-04 Asher J. Hancock , Allen Z. Ren , Anirudha Majumdar

Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich…

Robotics · Computer Science 2025-08-14 Lin Sun , Bin Xie , Yingfei Liu , Hao Shi , Tiancai Wang , Jiale Cao

End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zihao Sheng , Xin Ye , Jingru Luo , Sikai Chen , Liu Ren

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond…

Artificial Intelligence · Computer Science 2025-05-06 Yemin Shi , Yu Shu , Siwei Dong , Guangyi Liu , Jaward Sesay , Jingwen Li , Zhiting Hu

Current Vision-Language-Action (VLA) models are often constrained by a rigid, static interaction paradigm, which lacks the ability to see, hear, speak, and act concurrently as well as handle real-time user interruptions dynamically. This…

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models…

Robotics · Computer Science 2025-11-25 Tao Lin , Gen Li , Yilei Zhong , Yanwen Zou , Yuxin Du , Jiting Liu , Encheng Gu , Bo Zhao

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work is human-centered…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Dengming Zhang , Weitao You , Jingxiong Li , Weishen Lin , Wenda Shi , Xue Zhao , Heda Zuo , Junxian Wu , Lingyun Sun

Long-term action anticipation (LTA) aims to predict future actions over an extended period. Previous approaches primarily focus on learning exclusively from video data but lack prior knowledge. Recent researches leverage large language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Congqi Cao , Lanshu Hu , Yating Yu , Yanning Zhang

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autonomous driving tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yujin Wang , Quanfeng Liu , Zhengxin Jiang , Tianyi Wang , Junfeng Jiao , Hongqing Chu , Bingzhao Gao , Hong Chen