中文
相关论文

相关论文: Grounding World Simulation Models in a Real-World …

200 篇论文

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Shehan Munasinghe , Hanan Gani , Wenqi Zhu , Jiale Cao , Eric Xing , Fahad Shahbaz Khan , Salman Khan

World models represent a paradigm shift in generative AI, pursuing predictive understanding and controllable simulation of environments in a structured and generalizable way. We present World Machine, a generative world-modeling…

Learning robust robot policies in real-world environments requires diverse data augmentation, yet scaling real-world data collection is costly due to the need for acquiring physical assets and reconfiguring environments. Therefore,…

机器人学 · 计算机科学 2026-04-20 Jasper Lu , Zhenhao Shen , Yuanfei Wang , Shugao Liu , Shengqiang Xu , Shawn Xie , Jingkai Xu , Feng Jiang , Jade Yang , Chen Xie , Ruihai Wu

Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans,…

Video world models can generate realistic futures from a single instruction, but they often fail to preserve consistent point-level motion over time. As a result, the generated videos appear plausible, yet lack the physical grounding…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Kaichen Zhou , Yuzhen Chen , Fangneng Zhan , Hang Hua , Grace Chen , Xinhai Chang , Ao Qu , Yilun Du , Zhuang Liu , Paul Pu Liang , Mengyu Wang

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for visual entities.…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Xiangtai Li , Tao Zhang , Yanwei Li , Haobo Yuan , Shihao Chen , Yikang Zhou , Jiahao Meng , Yueyi Sun , Shilin Xu , Lu Qi , Tianheng Cheng , Yi Lin , Zilong Huang , Wenhao Huang , Jiashi Feng , Guang Shi

Foundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the…

机器人学 · 计算机科学 2025-07-18 Yucen Wang , Rui Yu , Shenghua Wan , Le Gan , De-Chuan Zhan

Data-efficient learning remains a central challenge in autonomous driving due to the high cost and safety risks of large-scale real-world interaction. Although world-model-based reinforcement learning enables policy optimization through…

机器人学 · 计算机科学 2026-03-10 Jiazhuo Li , Linjiang Cao , Qi Liu , Xi Xiong

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. Existing models, however, are typically restricted to limited state modalities, short video sequences, imprecise…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Bohan Li , Zhuang Ma , Dalong Du , Baorui Peng , Zhujin Liang , Zhenqiang Liu , Chao Ma , Yueming Jin , Hao Zhao , Wenjun Zeng , Xin Jin

We have seen great progress in basic perceptual tasks such as object recognition and detection. However, AI models still fail to match humans in high-level vision tasks due to the lack of capacities for deeper reasoning. Recently the new…

计算机视觉与模式识别 · 计算机科学 2016-04-12 Yuke Zhu , Oliver Groth , Michael Bernstein , Li Fei-Fei

We propose Infinite-World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ruiqi Wu , Xuanhua He , Meng Cheng , Tianyu Yang , Yong Zhang , Zhuoliang Kang , Xunliang Cai , Xiaoming Wei , Chunle Guo , Chongyi Li , Ming-Ming Cheng

The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limited diversity and fidelity with pronounced sim-to-real gaps.…

机器人学 · 计算机科学 2025-12-15 Junjie Ye , Rong Xue , Basile Van Hoorick , Pavel Tokmakov , Muhammad Zubair Irshad , Yue Wang , Vitor Guizilini

Generative world models offer a compelling foundation for augmented-reality (AR) applications: by predicting future image sequences that incorporate deliberate visual edits, they enable temporally coherent, augmented future frames that can…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Fanjun Bu , Chenyang Yuan , Hiroshi Yasuda

While Transformers have become the dominant architecture for visual generation, linear attention models, such as the state-space models (SSM), are increasingly recognized for their efficiency in processing long visual sequences. However,…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Yicong Hong , Long Mai , Yuan Yao , Feng Liu

This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods.…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Wenqiang Sun , Haiyu Zhang , Haoyuan Wang , Junta Wu , Zehan Wang , Zhenwei Wang , Yunhong Wang , Jun Zhang , Tengfei Wang , Chunchao Guo

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single…

Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated…

World models - learned internal simulators of environment dynamics - are rapidly becoming foundational to autonomous decision-making in robotics, autonomous vehicles, and agentic AI. By predicting future states in compressed latent spaces,…

密码学与安全 · 计算机科学 2026-04-08 Manoj Parmar

Accurate modeling and simulation of mobile networks are essential for enabling intelligent and cost-effective network optimization. In this paper, we propose MobiWorld, a generative world model designed to support high-fidelity and flexible…

网络与互联网体系结构 · 计算机科学 2025-07-15 Haoye Chai , Yuan Yuan , Yong Li

Deep vision models are now mature enough to be integrated in industrial and possibly critical applications such as autonomous navigation. Yet, data collection and labeling to train such models requires too much efforts and costs for a…

机器学习 · 计算机科学 2025-10-24 Estelle Chigot , Dennis G. Wilson , Meriem Ghrib , Fabrice Jimenez , Thomas Oberlin