基于对象中心体素化与神经渲染的动态场景理解
计算机视觉与模式识别
2025-02-17 v2
摘要
从无监督视频中学习对象中心表示具有挑战性。与大多数关注二维图像分解的先前方法不同,我们提出了一种名为 DynaVol-S 的 3D 生成模型,用于动态场景,使其在可微分体积渲染框架内实现对象中心学习。关键思想是对对象中心体素化,以捕捉场景的 3D 性质,这推断单个空间位置上的对象占用概率。这些体素特征通过标准空间变形函数演化,并在逆渲染管道中与组合性 NeRF 优化。此外,我们的方法将 2D 语义特征整合到 3D 语义网格中,通过多个解耦体素网格表示场景。DynaVol-S 在动态场景的新视角合成和无监督分解任务中显著优于现有模型。通过联合考虑几何结构和语义特征,它有效地解决了涉及复杂对象相互作用的挑战现实场景。 Furthermore, once trained, the explicitly meaningful voxel features enable additional capabilities that 2D scene decomposition methods cannot achieve, such as novel scene generation through editing geometric shapes or manipulating the motion trajectories of objects.
引用
@article{arxiv.2407.20908,
title = {Dynamic Scene Understanding through Object-Centric Voxelization and Neural Rendering},
author = {Yanpeng Zhao and Yiwei Hao and Siyu Gao and Yunbo Wang and Xiaokang Yang},
journal= {arXiv preprint arXiv:2407.20908},
year = {2025}
}
备注
Accepted by TPAMI2025