English

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

Computer Vision and Pattern Recognition 2026-06-30 v1

Abstract

The fundamental obstacle to industrial grade video generation is the lack of controllability: existing models treat video as a pixel distribution sampling problem, bypassing the explicit, instance level 4D4D (3D+T)(3D + T) physical world. Consequently, content creators cannot specify geometry, motion, camera parameters, or lighting in a deterministic, quantitative way, leading to the infamous ''gacha'' loop that makes professional content creation prohibitively inefficient and expensive. To address this, we introduce the World Narrative Model (WNM), a paradigm that decouples what to render -- the structured physical narrative -- from how to render -- the pixel generation process. WNM replaces end-to-end black-box sampling with orchestrated 4D4D pre-visualization for media generation. Collaborative agents translate sparse multimodal inputs, including text, reference videos, and sketches, into a fully editable world representation with scene geometry, object layouts, character/animal skeleton motion, trajectories, camera motion, and lighting at quantitative, physically meaningful granularity. This representation acts as a deterministic structural blueprint that drives existing video foundation models, either frozen or lightly adapted, to render final footage, turning the base model into a faithful neural shader. Built on this engine, our human-AI platform supports automatic world generation and pre-visualization aligned with professional filmmaking pipelines, while director consoles enable seamless human refinement. Experiments show that WNM greatly reduces probabilistic ``gacha'' calls and produces videos whose layout, motion, and cinematography closely follow creator intent. The framework is open and modular, allowing each component, such as world representation, control agents, and adapters, to be independently improved. Project website: https://glassroom.sjtu.edu.cn/WNM/.

Keywords

Cite

@article{arxiv.2606.31946,
  title  = {World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration},
  author = {Ye Chen and Xuanhong Chen and Yupeng Zhu and Liming Tan and Zhewen Wan and Yuxuan Xiong and Tielong Wang and Jinfan Liu and Wuze Zhang and Xiongzhen Zhang and Feifei Li and Xianglin Luo and Zhehan Zhao and Zhifan Zhang and Laisheng Kou and Zhujing Liang and Yugang Chen and Muchun Chen and Xu Miao and Yijing Zhang and Xiaojie Sheng and Qiang Hu and Jialiang Chen and Weimin Zhang and Wenjun Zhang and Bingbing Ni},
  journal= {arXiv preprint arXiv:2606.31946},
  year   = {2026}
}