English
Related papers

Related papers: DreamScene4D: Dynamic Multi-Object Scene Generatio…

200 papers

How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hanwen Liang , Junli Cao , Vidit Goel , Guocheng Qian , Sergei Korolev , Demetri Terzopoulos , Konstantinos N. Plataniotis , Sergey Tulyakov , Jian Ren

Recent advancements in generative models have enabled the creation of dynamic 4D content - 3D objects in motion - based on text prompts, which holds potential for applications in virtual worlds, media, and gaming. Existing methods provide…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Ohad Rahamim , Ori Malca , Dvir Samuel , Gal Chechik

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haozhe Xie , Zhaoxi Chen , Fangzhou Hong , Ziwei Liu

Recent advances in diffusion models have improved controllable streetscape generation and supported downstream perception and planning tasks. However, challenges remain in accurately modeling driving scenes and generating long videos. To…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jianbiao Mei , Tao Hu , Xuemeng Yang , Licheng Wen , Yu Yang , Tiantian Wei , Yukai Ma , Min Dou , Botian Shi , Yong Liu

We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available.…

Robotics · Computer Science 2018-02-27 Parv Parkhiya , Rishabh Khawad , J. Krishna Murthy , Brojeshwar Bhowmick , K. Madhava Krishna

We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Liao Shen , Xingyi Li , Huiqiang Sun , Juewen Peng , Ke Xian , Zhiguo Cao , Guosheng Lin

This paper presents a new method to synthesize an image from arbitrary views and times given a collection of images of a dynamic scene. A key challenge for the novel view synthesis arises from dynamic scene reconstruction where epipolar…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Jae Shin Yoon , Kihwan Kim , Orazio Gallo , Hyun Soo Park , Jan Kautz

We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Rafail Fridman , Amit Abecasis , Yoni Kasten , Tali Dekel

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

Scene extrapolation -- the idea of generating novel views by flying into a given image -- is a promising, yet challenging task. For each predicted frame, a joint inpainting and 3D refinement problem has to be solved, which is ill posed and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Shengqu Cai , Eric Ryan Chan , Songyou Peng , Mohamad Shahbazi , Anton Obukhov , Luc Van Gool , Gordon Wetzstein

This paper aims to tackle the challenge of dynamic view synthesis from multi-view videos. The key observation is that while previous grid-based methods offer consistent rendering, they fall short in capturing appearance details of a complex…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Haotong Lin , Sida Peng , Zhen Xu , Tao Xie , Xingyi He , Hujun Bao , Xiaowei Zhou

Text-to-3D generation has recently seen significant progress. To enhance its practicality in real-world applications, it is crucial to generate multiple independent objects with interactions, similar to layer-compositing in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Zizheng Yan , Jiapeng Zhou , Fanpeng Meng , Yushuang Wu , Lingteng Qiu , Zisheng Ye , Shuguang Cui , Guanying Chen , Xiaoguang Han

Recent advances in 3D scene generation produce visually appealing output, but current representations hinder artists' workflows that require modifiable 3D textured mesh scenes for visual effects and game development. Despite significant…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Tobias Sautter , Jan-Niklas Dihlmann , Hendrik P. A. Lensch

In this paper, we propose MoDGS, a new pipeline to render novel views of dy namic scenes from a casually captured monocular video. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid move ment of input…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Qingming Liu , Yuan Liu , Jiepeng Wang , Xianqiang Lyv , Peng Wang , Wenping Wang , Junhui Hou

To achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 In-Hwan Jin , Haesoo Choo , Seong-Hun Jeong , Heemoon Park , Junghwan Kim , Oh-joon Kwon , Kyeongbo Kong

Recovering 4D world from monocular video is a crucial yet challenging task. Conventional methods usually rely on the assumptions of multi-view videos, known camera parameters, or static scenes. In this paper, we relax all these constraints…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Shizun Wang , Xingyi Yang , Qiuhong Shen , Zhenxiang Jiang , Xinchao Wang

Generating explorable 3D scenes from a single image is a highly challenging problem in 3D vision. Existing methods struggle to support free exploration, often producing severe geometric distortions and noisy artifacts when the viewpoint…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Pengfei Wang , Liyi Chen , Zhiyuan Ma , Yanjun Guo , Guowen Zhang , Lei Zhang

This paper addresses the problem of decomposed 4D scene reconstruction from multi-view videos. Recent methods achieve this by lifting video segmentation results to a 4D representation through differentiable rendering techniques. Therefore,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Yongzhen Hu , Yihui Yang , Haotong Lin , Yifan Wang , Junting Dong , Yifu Deng , Xinyu Zhu , Fan Jia , Hujun Bao , Xiaowei Zhou , Sida Peng

When perceiving the world from multiple viewpoints, humans have the ability to reason about the complete objects in a compositional manner even when an object is completely occluded from certain viewpoints. Meanwhile, humans are able to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Chengmin Gao , Bin Li