English

ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

Computer Vision and Pattern Recognition 2025-10-08 v1

Abstract

Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single dynamic 3D representation end-to-end from the video. Our framework introduces three key components based on large-scale pre-trained 3D models: (i) a temporal attention that conditions generation on all frames while producing a time-indexed dynamic representation; (ii) a time-aware point sampling and 4D latent anchoring that promote temporally consistent geometry and texture; and (iii) noise sharing across frames to enhance temporal stability. Our method accurately captures non-rigid motion, volume changes, and even topological transitions without per-frame optimization. Across diverse in-the-wild videos, our method improves robustness and perceptual fidelity and reduces failure modes compared with the baselines.

Keywords

Cite

@article{arxiv.2510.06208,
  title  = {ShapeGen4D: Towards High Quality 4D Shape Generation from Videos},
  author = {Jiraphon Yenphraphai and Ashkan Mirzaei and Jianqi Chen and Jiaxu Zou and Sergey Tulyakov and Raymond A. Yeh and Peter Wonka and Chaoyang Wang},
  journal= {arXiv preprint arXiv:2510.06208},
  year   = {2025}
}

Comments

Project page: https://shapegen4d.github.io/

R2 v1 2026-07-01T06:22:06.326Z