English
Related papers

Related papers: Learning Primitive Embodied World Models: Towards …

200 papers

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language…

Machine Learning · Computer Science 2023-04-20 Yao Mu , Shunyu Yao , Mingyu Ding , Ping Luo , Chuang Gan

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and…

World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are high-dimensional and difficult to specify: for example, precisely controlling a human agent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Alex N. Wang , Trevor Darrell , Pavel Izmailov , Yutong Bai , Amir Bar

Vision-Language Models (VLMs) learn joint representations by mapping images and text into a shared latent space. However, recent research highlights that deterministic embeddings from standard VLMs often struggle to capture the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Aishwarya Venkataramanan , Paul Bodesheim , Joachim Denzler

The ability to perceive and reason about individual objects and their interactions is a goal to be achieved for building intelligent artificial systems. State-of-the-art approaches use a feedforward encoder to extract object embeddings and…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Jonathan Collu , Riccardo Majellaro , Aske Plaat , Thomas M. Moerland

Recent advances in vision, language, and multimodal learning have substantially accelerated progress in robotic foundation models, with robot manipulation remaining a central and challenging problem. This survey examines robot manipulation…

Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we…

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models…

Robotics · Computer Science 2026-04-06 Jiashu Yang , Yifan Han , Yucheng Xie , Ning Guo , Wenzhao Lian

Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap in standalone Vision-Language Models (VLMs), which often…

Artificial Intelligence · Computer Science 2026-05-08 Mohamed Salim Aissi , Clemence Grislain , Clement Romac , Laure Soulier , Mohamed Chetouani , Olivier Sigaud , Nicolas Thome

Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work trivially fitting the data without revealing the…

Robotics · Computer Science 2024-10-17 Kaidong Zhang , Pengzhen Ren , Bingqian Lin , Junfan Lin , Shikui Ma , Hang Xu , Xiaodan Liang

Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely underdeveloped. We introduce the Vision Language World Model…

Artificial Intelligence · Computer Science 2025-09-09 Delong Chen , Theo Moutakanni , Willy Chung , Yejin Bang , Ziwei Ji , Allen Bolourchi , Pascale Fung

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haotian Xue , Yipu Chen , Liqian Ma , Zelin Zhao , Lama Moukheiber , Yuchen Zhu , Yongxin Chen

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generation is computationally expensive and often unnecessary for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yiren Song , Yihan Wang , Xiyao Deng , Zhuoran Yan , Mike Zheng Shou

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge. While parameter-efficient prompt learning methods offer a promising path, they often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Enming Zhang , Jiayang Li , Yanru Wu , Zhenyu Liu , Yang Li

With hundreds of thousands of language models available on Huggingface today, efficiently evaluating and utilizing these models across various downstream, tasks has become increasingly critical. Many existing methods repeatedly learn…

Computation and Language · Computer Science 2024-10-18 Richard Zhuang , Tianhao Wu , Zhaojin Wen , Andrew Li , Jiantao Jiao , Kannan Ramchandran

What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificial environments by imagining all content. We present Seoul…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Junyoung Seo , Hyunwook Choi , Minkyung Kwon , Jinhyeok Choi , Siyoon Jin , Gayoung Lee , Junho Kim , JoungBin Lee , Geonmo Gu , Dongyoon Han , Sangdoo Yun , Seungryong Kim , Jin-Hwa Kim

In Model Predictive Control (MPC), world models predict the future outcomes of various action proposals, which are then scored to guide the selection of the optimal action. For visuomotor MPC, the score function is a distance metric between…

Robotics · Computer Science 2026-04-14 Quanyi Li , Lan Feng , Haonan Zhang , Wuyang Li , Letian Wang , Alexandre Alahi , Harold Soh

World models have gained significant attention as a promising approach for autonomous driving. By emulating human-like perception and decision-making processes, these models can predict and adapt to dynamic environments. Existing methods…

Robotics · Computer Science 2025-12-03 Huiqian Li , Wei Pan , Haodong Zhang , Jin Huang , Zhihua Zhong