中文
相关论文

相关论文: ConsDreamer: Advancing Multi-View Consistency for …

200 篇论文

This paper introduces a novel approach to synthesize texture to dress up a given 3D object, given a text prompt. Based on the pretrained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-inpaint approach, in…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Yuxin Liu , Minshan Xie , Hanyuan Liu , Tien-Tsin Wong

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to accommodate…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Jiahui Yang , Donglin Di , Baorui Ma , Xun Yang , Yongjia Ma , Wenzhang Sun , Wei Chen , Jianxun Cui , Zhou Xue , Meng Wang , Yebin Liu

Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Yuanxun Lu , Jingyang Zhang , Shiwei Li , Tian Fang , David McKinnon , Yanghai Tsin , Long Quan , Xun Cao , Yao Yao

Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Zheyuan Liu , Munan Ning , Qihui Zhang , Shuo Yang , Zhongrui Wang , Yiwei Yang , Xianzhe Xu , Yibing Song , Weihua Chen , Fan Wang , Li Yuan

While text-to-3D and image-to-3D generation tasks have received considerable attention, one important but under-explored field between them is controllable text-to-3D generation, which we mainly focus on in this work. To address this task,…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Zhiqi Li , Yiming Chen , Lingzhe Zhao , Peidong Liu

Synthesizing consistent and photorealistic 3D scenes is an open problem in computer vision. Video diffusion models generate impressive videos but cannot directly synthesize 3D representations, i.e., lack 3D consistency in the generated…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Katja Schwarz , Norman Mueller , Peter Kontschieder

We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Kyle Sargent , Zizhang Li , Tanmay Shah , Charles Herrmann , Hong-Xing Yu , Yunzhi Zhang , Eric Ryan Chan , Dmitry Lagun , Li Fei-Fei , Deqing Sun , Jiajun Wu

Recently, the field of text-guided 3D scene generation has garnered significant attention. High-quality generation that aligns with physical realism and high controllability is crucial for practical 3D scene applications. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Yang Zhou , Zongjin He , Qixuan Li , Chao Wang

Recent one image to 3D generation methods commonly adopt Score Distillation Sampling (SDS). Despite the impressive results, there are multiple deficiencies including multi-view inconsistency, over-saturated and over-smoothed textures, as…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Junwu Zhang , Zhenyu Tang , Yatian Pang , Xinhua Cheng , Peng Jin , Yida Wei , Munan Ning , Li Yuan

Score Distillation Sampling (SDS) has emerged as a prominent method for text-to-3D generation by leveraging the strengths of 2D diffusion models. However, SDS is limited to generation tasks and lacks the capability to edit existing 3D…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Xingyu Miao , Haoran Duan , Yang Long , Jungong Han

Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. In this work, we propose \textbf{ViewCrafter}, a novel method for synthesizing high-fidelity novel…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Wangbo Yu , Jinbo Xing , Li Yuan , Wenbo Hu , Xiaoyu Li , Zhipeng Huang , Xiangjun Gao , Tien-Tsin Wong , Ying Shan , Yonghong Tian

Recent CLIP-guided 3D optimization methods, such as DreamFields and PureCLIPNeRF, have achieved impressive results in zero-shot text-to-3D synthesis. However, due to scratch training and random initialization without prior knowledge, these…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Jiale Xu , Xintao Wang , Weihao Cheng , Yan-Pei Cao , Ying Shan , Xiaohu Qie , Shenghua Gao

Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Josef Bengtson , David Nilsson , Dong In Lee , Yaroslava Lochman , Fredrik Kahl

Novel view synthesis under sparse views has been a long-term important challenge in 3D reconstruction. Existing works mainly rely on introducing external semantic or depth priors to supervise the optimization of 3D representations. However,…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Qisen Wang , Yifan Zhao , Jiawei Ma , Jia Li

Score distillation sampling (SDS) and its variants have greatly boosted the development of text-to-3D generation, but are vulnerable to geometry collapse and poor textures yet. To solve this issue, we first deeply analyze the SDS and find…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Zike Wu , Pan Zhou , Xuanyu Yi , Xiaoding Yuan , Hanwang Zhang

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Hyojun Go , Byeongjun Park , Hyelin Nam , Byung-Hoon Kim , Hyungjin Chung , Changick Kim

3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, the…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Haiyang Zhou , Xinhua Cheng , Wangbo Yu , Yonghong Tian , Li Yuan

Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models are limited by…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Ye Tao , Jiawei Zhang , Yahao Shi , Dongqing Zou , Bin Zhou

3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times, some work has emerged that aims to extend…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Shijie Zhou , Haoran Chang , Sicheng Jiang , Zhiwen Fan , Zehao Zhu , Dejia Xu , Pradyumna Chari , Suya You , Zhangyang Wang , Achuta Kadambi