English
Related papers

Related papers: Scene Co-pilot: Procedural Text to Video Generatio…

200 papers

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propose a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Fei Yu , Quan Deng , Shengeng Tang , Yuehua Li , Lechao Cheng

3D scene generation plays a crucial role in gaming, artistic creation, virtual reality, and many other domains. However, current 3D scene design still relies heavily on extensive manual effort from creators, and existing automated methods…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Zixuan Bian , Ruohan Ren , Yue Yang , Chris Callison-Burch

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Zhi Cen , Huaijin Pi , Sida Peng , Zehong Shen , Minghui Yang , Shuai Zhu , Hujun Bao , Xiaowei Zhou

3D scene generation seeks to synthesize spatially structured, semantically meaningful, and photorealistic environments for applications such as immersive media, robotics, autonomous driving, and embodied AI. Early methods based on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Beichen Wen , Haozhe Xie , Zhaoxi Chen , Fangzhou Hong , Ziwei Liu

3D content creation has long been a complex and time-consuming process, often requiring specialized skills and resources. While recent advancements have allowed for text-guided 3D object and scene generation, they still fall short of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Xingyi Li , Yizheng Wu , Jun Cen , Juewen Peng , Kewei Wang , Ke Xian , Zhe Wang , Zhiguo Cao , Guosheng Lin

Generating 3D scenes from natural language holds great promise for applications in gaming, film, and design. However, existing methods struggle with automation, 3D consistency, and fine-grained control. We present DreamScene, an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Haoran Li , Yuli Tian , Kun Lan , Yong Liao , Lin Wang , Pan Hui , Peng Yuan Zhou

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

Creating music is iterative, requiring varied methods at each stage. However, existing AI music systems fall short in orchestrating multiple subsystems for diverse needs. To address this gap, we introduce Loop Copilot, a novel system that…

Sound · Computer Science 2024-09-02 Yixiao Zhang , Akira Maezawa , Gus Xia , Kazuhiko Yamamoto , Simon Dixon

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still exhibit certain shortcomings,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Shuangkang Fang , Yufeng Wang , Yi-Hsuan Tsai , Yi Yang , Wenrui Ding , Shuchang Zhou , Ming-Hsuan Yang

We introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Seok-Young Kim , Dooyoung Kim , Woojin Cho , Hail Song , Suji Kang , Woontack Woo

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Minjun Kang , Inkyu Shin , Taeyeop Lee , In So Kweon , Kuk-Jin Yoon

We present a system for generating indoor scenes in response to text prompts. The prompts are not limited to a fixed vocabulary of scene descriptions, and the objects in generated scenes are not restricted to a fixed set of object…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Rio Aguina-Kang , Maxim Gumin , Do Heon Han , Stewart Morris , Seung Jean Yoo , Aditya Ganeshan , R. Kenny Jones , Qiuhong Anna Wei , Kailiang Fu , Daniel Ritchie

In the field of 3D content generation, single image scene reconstruction methods still struggle to simultaneously ensure the quality of individual assets and the coherence of the overall scene in complex environments, while texture editing…

Graphics · Computer Science 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Qinglin Zeng , Kaitong Cai , Ruiqi Chen , Qinhan Lv , Keze Wang

Perpetual 3D scene generation aims to produce long-range and coherent 3D view sequences, which is applicable for long-term video synthesis and 3D scene reconstruction. Existing methods follow a "navigate-and-imagine" fashion and rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Chong Xia , Shengjun Zhang , Fangfu Liu , Chang Liu , Khodchaphun Hirunyaratsameewong , Yueqi Duan

This paper presents a novel generative approach that outputs 3D indoor environments solely from a textual description of the scene. Current methods often treat scene synthesis as a mere layout prediction task, leading to rooms with…

Machine Learning · Computer Science 2025-02-12 Yao Wei , Matteo Toso , Pietro Morerio , Michael Ying Yang , Alessio Del Bue

We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Rafail Fridman , Amit Abecasis , Yoni Kasten , Tali Dekel

We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Quynh Phung , Long Mai , Fabian David Caba Heilbron , Feng Liu , Jia-Bin Huang , Cusuh Ham

Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Ziyan Guo , Haoxuan Qu , Hossein Rahmani , Dewen Soh , Ping Hu , Qiuhong Ke , Jun Liu

Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Gonzalo Gomez-Nogales , Yicong Hong , Chongjian Ge , Peiye Zhuang , Marc Comino-Trinidad , Dan Casas , Yi Zhou