English
Related papers

Related papers: Advancing Open-source World Models

200 papers

Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Zangwei Zheng , Xiangyu Peng , Tianji Yang , Chenhui Shen , Shenggui Li , Hongxin Liu , Yukun Zhou , Tianyi Li , Yang You

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Xiaofeng Mao , Zhen Li , Chuanhao Li , Xiaojie Xu , Kaining Ying , Tong He , Jiangmiao Pang , Yu Qiao , Kaipeng Zhang

We introduce LongCat-Image, a pioneering open-source and bilingual (Chinese-English) foundation model for image generation, designed to address core challenges in multilingual text rendering, photorealism, deployment efficiency, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Meituan LongCat Team , Hanghang Ma , Haoxian Tan , Jiale Huang , Junqiang Wu , Jun-Yan He , Lishuai Gao , Songlin Xiao , Xiaoming Wei , Xiaoqi Ma , Xunliang Cai , Yayong Guan , Jie Hu

While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building agents that can survive and thrive in the real world (for…

Web agents require massive trajectories to generalize, yet real-world training is constrained by network latency, rate limits, and safety risks. We introduce \textbf{WebWorld} series, the first open-web simulator trained at scale. While…

Artificial Intelligence · Computer Science 2026-02-17 Zikai Xiao , Jianhong Tu , Chuhang Zou , Yuxin Zuo , Zhi Li , Peng Wang , Bowen Yu , Fei Huang , Junyang Lin , Zuozhu Liu

We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Hyeongju Mun , In-Hwan Jin , Sohyeong Kim , Kyeongbo Kong

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Video generation models (VGMs) have received extensive attention recently and serve as promising candidates for general-purpose large vision models. While they can only generate short videos each time, existing methods achieve long video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yuanhui Huang , Wenzhao Zheng , Yuan Gao , Xin Tao , Pengfei Wan , Di Zhang , Jie Zhou , Jiwen Lu

Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user control over the environment for reproducible, editable…

Artificial Intelligence · Computer Science 2026-04-01 Ryan Po , David Junhao Zhang , Amir Hertz , Gordon Wetzstein , Neal Wadhwa , Nataniel Ruiz

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While recent video generation models produce realistic visual…

Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Samuel Garcin , Thomas Walker , Steven McDonagh , Tim Pearce , Hakan Bilen , Tianyu He , Kaixin Wang , Jiang Bian

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jianxiong Gao , Zhaoxi Chen , Xian Liu , Junhao Zhuang , Chengming Xu , Jianfeng Feng , Yu Qiao , Yanwei Fu , Chenyang Si , Ziwei Liu

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical…

Image and Video Processing · Electrical Eng. & Systems 2026-05-06 Muyang He , Hanzhong Guo , Junxiong Lin , Yizhou Yu

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects in isolation, as…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yicong Hong , Yiqun Mei , Chongjian Ge , Yiran Xu , Yang Zhou , Sai Bi , Yannick Hold-Geoffroy , Mike Roberts , Matthew Fisher , Eli Shechtman , Kalyan Sunkavalli , Feng Liu , Zhengqi Li , Hao Tan

World simulators can provide safe and scalable environments for training Physical AI systems before real-world deployment. Large video generation models are emerging as a promising basis for such simulators because they can generate diverse…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Pu Zhao , Juyi Lin , Timothy Rupprecht , Arash Akbari , Chence Yang , Rahul Chowdhury , Elaheh Motamedi , Arman Akbari , Yumei He , Chen Wang , Geng Yuan , Weiwei Chen , Yanzhi Wang

Large Language Models (LLMs) have proven their worth across a diverse spectrum of disciplines. LLMs have shown great potential in Procedural Content Generation (PCG) as well, but directly generating a level through a pre-trained LLM is…

Computation and Language · Computer Science 2024-05-14 Muhammad U. Nasir , Steven James , Julian Togelius

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zhiheng Liu , Xueqing Deng , Shoufa Chen , Angtian Wang , Qiushan Guo , Mingfei Han , Zeyue Xue , Mengzhao Chen , Ping Luo , Linjie Yang

World models are central to building agents that can reason, plan, and generalize beyond their training data. However, research on world models is currently fragmented, with disparate codebases, data pipelines, and evaluation protocols…

Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for…