中文

生成式相机推车:极端单目动态新视角合成

计算机视觉与模式识别 2024-07-08 v2 人工智能 机器学习 机器人学

摘要

从单一视角准确重建复杂动态场景在计算机视觉中仍是具有挑战性的任务。当前的动态新视角合成方法通常需要来自多个不同相机视角的视频, necessitating careful recording setups, significantly restricting their utility in the wild as well as in terms of embodied AI applications. In this paper, we propose GCD\textbf{GCD}, a controllable monocular dynamic view synthesis pipeline that leverages large-scale diffusion priors to, given a video of any scene, generate a synchronous video from any other chosen perspective, conditioned on a set of relative camera pose parameters. Our model does not require depth as input, and does not explicitly model 3D scene geometry, instead performing end-to-end video-to-video translation in order to achieve its goal efficiently. Despite being trained on synthetic multi-view video data only, zero-shot real-world generalization experiments show promising results in multiple domains, including robotics, object permanence, and driving environments. We believe our framework can potentially unlock powerful applications in rich dynamic scene understanding, perception for robotics, and interactive 3D video viewing experiences for virtual reality.

关键词

引用

@article{arxiv.2405.14868,
  title  = {Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis},
  author = {Basile Van Hoorick and Rundi Wu and Ege Ozguroglu and Kyle Sargent and Ruoshi Liu and Pavel Tokmakov and Achal Dave and Changxi Zheng and Carl Vondrick},
  journal= {arXiv preprint arXiv:2405.14868},
  year   = {2024}
}

备注

Accepted to ECCV 2024. Project webpage is available at: https://gcd.cs.columbia.edu/