中文
相关论文

相关论文: Bootstrapping Action-Grounded Visual Dynamics in U…

200 篇论文

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often leading to…

机器人学 · 计算机科学 2025-09-10 Qi Lv , Weijie Kong , Hao Li , Jia Zeng , Zherui Qiu , Delin Qu , Haoming Song , Qizhi Chen , Xiang Deng , Jiangmiao Pang

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geometry, effectively anticipating the future requires capturing…

机器人学 · 计算机科学 2026-03-12 Xiaoxu Xu , Hao Li , Jinhui Ye , Yilun Chen , Jia Zeng , Xinyi Chen , Linning Xu , Dahua Lin , Weixin Li , Jiangmiao Pang

Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text training data. Curation of…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Giorgio Giannone , Ruoteng Li , Qianli Feng , Evgeny Perevodchikov , Rui Chen , Aleix Martinez

Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce…

机器人学 · 计算机科学 2025-09-29 Asher J. Hancock , Xindi Wu , Lihan Zha , Olga Russakovsky , Anirudha Majumdar

Multimodal learning plays a critical role in e-commerce recommendation platforms today, enabling accurate recommendations and product understanding. However, existing vision-language models, such as CLIP, face key challenges in e-commerce…

A reliable driving assistant should provide consistent responses based on temporally grounded reasoning derived from observed information. In this work, we investigate whether Vision-Language Models (VLMs), when applied as driving…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Chun-Peng Chang , Chen-Yu Wang , Holger Caesar , Alain Pagani

World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision-language models (VLMs) for planning, yet existing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zhexiao Xiong , Xin Ye , Burhan Yaman , Sheng Cheng , Yiren Lu , Jingru Luo , Nathan Jacobs , Liu Ren

Imitation learning is a prominent paradigm for robotic manipulation. However, existing visual imitation methods map 2D image observations directly to 3D action outputs, imposing a 2D-3D mismatch that hinders spatial reasoning and degrades…

机器人学 · 计算机科学 2026-03-19 Tianxing Zhou , Feiyang Xue , Zhangchen Ye , Tianyuan Yuan , Hang Zhao , Tao Jiang

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Shijie Zhou , Alexander Vilesov , Xuehai He , Ziyu Wan , Shuwang Zhang , Aditya Nagachandra , Di Chang , Dongdong Chen , Xin Eric Wang , Achuta Kadambi

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

Foundation models (FMs) have shown prominent success in a wide range of tasks. Their applicability to specific domain-task pairings relies on the availability of, both, high-quality data and significant computational resources. These…

机器学习 · 计算机科学 2023-11-10 Lekang Jiang , Filip Svoboda , Nicholas D. Lane

Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing…

机器学习 · 计算机科学 2025-02-04 Shiqi He , Insu Jang , Mosharaf Chowdhury

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jiangtao Wu , Shihao Li , Zhaozhou Bian , Jialu Chen , Runzhe Wen , An Ping , Yiwen He , Jiakai Wang , Yuanxing Zhang , Jiaheng Liu

Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GPT-5 still struggle in understanding 3D from 2D inputs. On…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zhipeng Cai , Ching-Feng Yeh , Hu Xu , Zhuang Liu , Gregory Meyer , Xinjie Lei , Changsheng Zhao , Shang-Wen Li , Vikas Chandra , Yangyang Shi

Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in representation learning. Due to image-level visual-language…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Mengxiao Tian , Xinxiao Wu , Shuo Yang

Autonomous driving has seen significant progress, driven by extensive real-world data. However, in long-tail scenarios, accurately predicting the safety of the ego vehicle's future motion remains a major challenge due to uncertainties in…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Zhiyi Hou , Enhui Ma , Fang Li , Zhiyi Lai , Kalok Ho , Zhanqian Wu , Lijun Zhou , Long Chen , Chitian Sun , Haiyang Sun , Bing Wang , Guang Chen , Hangjun Ye , Kaicheng Yu

Vision-Language Pre-training (VLP) aims to learn multi-modal representations from image-text pairs and serves for downstream vision-language tasks in a fine-tuning fashion. The dominant VLP models adopt a CNN-Transformer architecture, which…

计算机视觉与模式识别 · 计算机科学 2021-11-10 Hongwei Xue , Yupan Huang , Bei Liu , Houwen Peng , Jianlong Fu , Houqiang Li , Jiebo Luo

Ensuring safe navigation in complex environments requires accurate real-time traversability assessment and understanding of environmental interactions relative to the robot`s capabilities. Traditional methods, which assume simplified…

机器人学 · 计算机科学 2025-04-30 Pascal Roth , Jonas Frey , Cesar Cadena , Marco Hutter

Although Multimodal Large Language Models (MLLMs) have demonstrated promising versatile capabilities, their performance is still inferior to specialized models on downstream tasks, which makes adaptation necessary to enhance their utility.…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Yichi Zhang , Yinpeng Dong , Siyuan Zhang , Tianzan Min , Hang Su , Jun Zhu

Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as counting or relational reasoning. To understand the…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Zihan Weng , Lucas Gomez , Taylor Whittington Webb , Pouya Bashivan