English

SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation

Computer Vision and Pattern Recognition 2026-03-31 v2

Abstract

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach that learns the continuous 4D dynamics and generate the unseen visual contents. The principle behind SeeU is a new 2D\to4D\to2D learning framework. SeeU first reconstructs the 4D world from sparse and monocular 2D frames (2D\to4D). It then learns the continuous 4D dynamics on a low-rank representation and physical constraints (discrete 4D\tocontinuous 4D). Finally, SeeU rolls the world forward in time, re-projects it back to 2D at sampled times and viewpoints, and generates unseen regions based on spatial-temporal context awareness (4D\to2D). By modeling dynamics in 4D, SeeU achieves continuous and physically-consistent novel visual generation, demonstrating strong potentials in multiple tasks including unseen temporal generation, unseen spatial generation, and video editing. All data and code will be public at https://yuyuanspace.com/SeeU/

Keywords

Cite

@article{arxiv.2512.03350,
  title  = {SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation},
  author = {Yu Yuan and Tharindu Wickremasinghe and Zeeshan Nadir and Xijun Wang and Yiheng Chi and Stanley H. Chan},
  journal= {arXiv preprint arXiv:2512.03350},
  year   = {2026}
}

Comments

Accepted by CVPR 2026. Camera-Ready Version. Project Page: https://yuyuanspace.com/SeeU/

R2 v1 2026-07-01T08:06:53.060Z