English
Related papers

Related papers: PhiP-G: Physics-Guided Text-to-3D Compositional Sc…

200 papers

Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -video approaches struggle to maintain scene- and object-level consistency beyond a limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Manuel-Andreas Schneider , Angela Dai

Despite the recent impressive breakthroughs in text-to-image generation, generative models have difficulty in capturing the data distribution of underrepresented attribute compositions while over-memorizing overrepresented attribute…

Computer Vision and Pattern Recognition · Computer Science 2023-01-05 Yuren Cong , Martin Renqiang Min , Li Erran Li , Bodo Rosenhahn , Michael Ying Yang

3D Gaussian Splatting (3DGS) has garnered significant attention due to its superior scene representation fidelity and real-time rendering performance, especially for dynamic 3D scene reconstruction (\textit{i.e.}, 4D reconstruction).…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Henan Wang , Hanxin Zhu , Xinliang Gong , Tianyu He , Xin Li , Zhibo Chen

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based methods improve object arrangements using spatial constraints…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Zeeshan Khan , Shizhe Chen , Cordelia Schmid

Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Yunhan Yang , Chunshi Wang , Junliang Ye , Yang Li , Zanxin Chen , Zehuan Huang , Yao Mu , Zhuo Chen , Chunchao Guo , Xihui Liu

Despite significant advancements in text-to-image models for generating high-quality images, these methods still struggle to ensure the controllability of text prompts over images in the context of complex text prompts, especially when it…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Zhenyu Wang , Enze Xie , Aoxue Li , Zhongdao Wang , Xihui Liu , Zhenguo Li

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle…

Computer Vision and Pattern Recognition · Computer Science 2021-04-30 Michael Niemeyer , Andreas Geiger

Recent 3D generative models can synthesize high-quality assets, but their outputs are typically static: they lack the skeletal rigs, joint hierarchies, and skinning weights required for animation. This limits their use in games, film,…

Graphics · Computer Science 2026-05-14 Nikitas Chatzis , Marios Loizou , Evangelos Kalogerakis

Despite their impressive realism, modern text-to-image models still struggle with compositionality, often failing to render accurate object counts, attributes, and spatial relations. To address this challenge, we present a training-free…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Minsuk Ji , Sanghyeok Lee , Namhyuk Ahn

Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Amir Mohammad Izadi , Seyed Mohammad Hadi Hosseini , Soroush Vafaie Tabar , Ali Abdollahi , Armin Saghafian , Mahdieh Soleymani Baghshah

3D modeling is moving from virtual to physical. Existing 3D generation primarily emphasizes geometries and textures while neglecting physical-grounded modeling. Consequently, despite the rapid development of 3D generative models, the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Ziang Cao , Zhaoxi Chen , Liang Pan , Ziwei Liu

3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually improved, the performance of existing works is still far…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Yuanyuan Liu , Chengjiang Long , Zhaoxuan Zhang , Bokai Liu , Qiang Zhang , Baocai Yin , Xin Yang

Creating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Yian Wang , Xiaowen Qiu , Jiageng Liu , Zhehuan Chen , Jiting Cai , Yufei Wang , Tsun-Hsuan Wang , Zhou Xian , Chuang Gan

Synthesizing 3D scenes from open-vocabulary text descriptions is a challenging, important, and recently-popular application. One of its critical subproblems is layout generation: given a set of objects, lay them out to produce a scene…

Developing comprehensive explicit world models is crucial for understanding and simulating real-world scenarios. Recently, Procedural Controllable Generation (PCG) has gained significant attention in large-scale scene generation by enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Mengqi Zhou , Yuxi Wang , Jun Hou , Shougao Zhang , Yiwei Li , Chuanchen Luo , Junran Peng , Zhaoxiang Zhang

Compositional scene reconstruction seeks to create object-centric representations rather than holistic scenes from real-world videos, which is natively applicable for simulation and interaction. Conventional compositional reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Chong Xia , Kai Zhu , Zizhuo Wang , Fangfu Liu , Zhizheng Zhang , Yueqi Duan

Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Lu Ling , Chen-Hsuan Lin , Tsung-Yi Lin , Yifan Ding , Yu Zeng , Yichen Sheng , Yunhao Ge , Ming-Yu Liu , Aniket Bera , Zhaoshuo Li

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Hyojun Go , Byeongjun Park , Hyelin Nam , Byung-Hoon Kim , Hyungjin Chung , Changick Kim

Recent advances in text-to-vision generation excel in visual fidelity but struggle with compositional generalization and semantic alignment. Existing datasets are noisy and weakly compositional, limiting models' understanding of complex…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Ziqi Gao , Weikai Huang , Jieyu Zhang , Aniruddha Kembhavi , Ranjay Krishna
‹ Prev 1 4 5 6 7 8 10 Next ›