English
Related papers

Related papers: CoCo4D: Comprehensive and Complex 4D Scene Generat…

200 papers

Automatic scene generation is an essential area of research with applications in robotics, recreation, visual representation, training and simulation, education, and more. This survey provides a comprehensive review of the current…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Awal Ahmed Fime , Saifuddin Mahmud , Arpita Das , Md. Sunzidul Islam , Hong-Hoon Kim

Generative models have shown substantial impact across multiple domains, their potential for scene synthesis remains underexplored in robotics. This gap is more evident in drone simulators, where simulation environments still rely heavily…

Controllable scene synthesis aims to create interactive environments for various industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applications by abstracting the scene context in a compact manner.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Guangyao Zhai , Evin Pınar Örnek , Shun-Cheng Wu , Yan Di , Federico Tombari , Nassir Navab , Benjamin Busam

This report surveys advances in deep learning-based modeling techniques that address four different 3D indoor scene analysis tasks, as well as synthesis of 3D indoor scenes. We describe different kinds of representations for indoor scenes,…

Graphics · Computer Science 2023-08-22 Akshay Gadi Patil , Supriya Gadi Patil , Manyi Li , Matthew Fisher , Manolis Savva , Hao Zhang

Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Gonzalo Gomez-Nogales , Yicong Hong , Chongjian Ge , Peiye Zhuang , Marc Comino-Trinidad , Dan Casas , Yi Zhou

The vision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, the presence of dynamic objects in the scene seriously affects the accuracy of the model…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Meng Wang , Fan Wu , Yunchuan Qin , Ruihui Li , Zhuo Tang , Kenli Li

Instruction-guided generative models, especially those using text-to-image (T2I) and text-to-video (T2V) diffusion frameworks, have advanced the field of content editing in recent years. To extend these capabilities to 4D scene, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hasan Iqbal , Nazmul Karim , Umar Khalid , Azib Farooq , Zichun Zhong , Chen Chen , Jing Hua

Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of synthesizing temporally coherent dynamic 3D assets guided…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Qiaowei Miao , Kehan Li , Jinsheng Quan , Zhiyuan Min , Shaojie Ma , Yichao Xu , Yi Yang , Ping Liu , Yawei Luo

Automatically generating a complete 3D scene from a text description, a reference image, or both has significant applications in fields like virtual reality and gaming. However, current methods often generate low-quality textures and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Zhexiao Xiong , Zhang Chen , Zhong Li , Yi Xu , Nathan Jacobs

Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Kaichen Zhou , Yuhan Wang , Grace Chen , Xinhai Chang , Gaspard Beaudouin , Fangneng Zhan , Paul Pu Liang , Mengyu Wang

Recent 4D dynamic scene editing methods require editing thousands of 2D images used for dynamic scene synthesis and updating the entire scene with additional training loops, resulting in several hours of processing to edit a single dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Joohyun Kwon , Hanbyel Cho , Junmo Kim

Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Jiraphon Yenphraphai , Jianqi Chen , Jian Wang , Gordon Qian , Sergey Tulyakov , Rameen Abdal , Raymond A. Yeh , Peter Wonka , Chaoyang Wang

Recent text-to-image generation methods provide a simple yet exciting conversion capability between text and image domains. While these methods have incrementally improved the generated image fidelity and text relevancy, several pivotal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Oran Gafni , Adam Polyak , Oron Ashual , Shelly Sheynin , Devi Parikh , Yaniv Taigman

We introduce a method to synthesize animator guided human motion across 3D scenes. Given a set of sparse (3 or 4) joint locations (such as the location of a person's hand and two feet) and a seed motion sequence in a 3D scene, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Aymen Mir , Xavier Puig , Angjoo Kanazawa , Gerard Pons-Moll

Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Thomas Tanay , Mohammed Brahimi , Michal Nazarczuk , Qingwen Zhang , Sibi Catley-Chandar , Arthur Moreau , Zhensong Zhang , Eduardo Pérez-Pellitero

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiazhe Guo , Yikang Ding , Xiwu Chen , Shuo Chen , Bohan Li , Yingshuang Zou , Xiaoyang Lyu , Feiyang Tan , Xiaojuan Qi , Zhiheng Li , Hao Zhao

Text-driven 3D scene generation is widely applicable to video gaming, film industry, and metaverse applications that have a large demand for 3D scenes. However, existing text-to-3D generation methods are limited to producing 3D objects with…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Jingbo Zhang , Xiaoyu Li , Ziyu Wan , Can Wang , Jing Liao

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Hongjie Li , Hong-Xing Yu , Jiaman Li , Jiajun Wu

Despite recent advances in text-conditioned 3D indoor scene generation, there remain gaps in the evaluation of these methods. Existing metrics often measure realism by comparing generated scenes to a set of ground-truth scenes, but they…

Graphics · Computer Science 2026-03-10 Hou In Ivan Tam , Hou In Derek Pun , Austin T. Wang , Angel X. Chang , Manolis Savva

Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Pinxin Liu , Pengfei Zhang , Hyeongwoo Kim , Pablo Garrido , Ari Shapiro , Kyle Olszewski