We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified modeling of appearance, geometry and motion from monocular videos, and enables reconstruction, view synthesis and 3D point tracking within a single learning-based framework. By bridging view synthesis with geometry reconstruction, MoVieS enables large-scale training on diverse datasets with minimal dependence on task-specific supervision. As a result, it also naturally supports a wide range of zero-shot applications, such as scene flow estimation and moving object segmentation. Extensive experiments validate the effectiveness and efficiency of MoVieS across multiple tasks, achieving competitive performance while offering several orders of magnitude speedups.
@article{arxiv.2507.10065,
title = {MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second},
author = {Chenguo Lin and Yuchen Lin and Panwang Pan and Yifan Yu and Tao Hu and Honglei Yan and Katerina Fragkiadaki and Yadong Mu},
journal= {arXiv preprint arXiv:2507.10065},
year = {2026}
}
Comments
Project page: https://chenguolin.github.io/projects/MoVieS; Accepted to CVPR 2026