English

SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splatting

Computer Vision and Pattern Recognition 2026-05-07 v2

Abstract

Reconstructing a dynamic target moving over a large area is challenging. Standard approaches for dynamic object reconstruction require dense coverage in both the viewing space and the temporal dimension, typically relying on multi-view videos captured at each time step. However, such setups are only possible in constrained environments. In real-world scenarios, observations are often sparse over time and captured sparsely from diverse viewpoints (e.g., from security cameras), making dynamic reconstruction highly ill-posed. We present SV-GS, a framework that simultaneously estimates a deformation model and the object's motion over time under sparse observations. To initialize SV-GS, we leverage a rough skeleton graph and an initial static reconstruction as inputs to guide motion estimation. (Later, we show that this input requirement can be relaxed.) Our method optimizes a skeleton-driven deformation field composed of a coarse skeleton joint pose estimator and a module for fine-grained deformations. By making only the joint pose estimator time-dependent, our model enables smooth motion interpolation while preserving learned geometric details. Experiments on synthetic datasets show that our method outperforms existing approaches under sparse observations by up to 34% in PSNR, and achieves comparable performance to dense monocular video methods on real-world datasets despite using significantly fewer frames. Moreover, we demonstrate that the input initial static reconstruction can be replaced by a diffusion-based generative prior, making our method more practical for real-world scenarios.

Keywords

Cite

@article{arxiv.2601.00285,
  title  = {SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splatting},
  author = {Jun-Jee Chao and Volkan Isler},
  journal= {arXiv preprint arXiv:2601.00285},
  year   = {2026}
}
R2 v1 2026-07-01T08:47:45.398Z