English
Related papers

Related papers: Taming Mode Collapse in Score Distillation for Tex…

200 papers

This work presents HeadArtist for 3D head generation from text descriptions. With a landmark-guided ControlNet serving as the generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Hongyu Liu , Xuan Wang , Ziyu Wan , Yujun Shen , Yibing Song , Jing Liao , Qifeng Chen

We propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function implicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Xinyu Peng , Ziyang Zheng , Yaoming Wang , Han Li , Nuowen Kan , Wenrui Dai , Chenglin Li , Junni Zou , Hongkai Xiong

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Chengyue Wu , Xiaokang Chen , Zhiyu Wu , Yiyang Ma , Xingchao Liu , Zizheng Pan , Wen Liu , Zhenda Xie , Xingkai Yu , Chong Ruan , Ping Luo

Dataset distillation aims to synthesize a small dataset from a large dataset, enabling the model trained on it to perform well on the original dataset. With the blooming of large language models and multimodal large language models, the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Zhenghao Zhao , Haoxuan Wang , Junyi Wu , Yuzhang Shang , Gaowen Liu , Yan Yan

Dataset distillation aims to synthesize a compact dataset from the original large-scale one, enabling highly efficient learning while preserving competitive model performance. However, traditional techniques primarily capture low-level…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Qianxin Xia , Jiawei Du , Guoming Lu , Zhiyong Shu , Jielei Wang

In this paper, we propose an effective two-stage approach named Grounded-Dreamer to generate 3D assets that can accurately follow complex, compositional text prompts while achieving high fidelity by using a pre-trained multi-view diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Xiaolong Li , Jiawei Mo , Ying Wang , Chethan Parameshwara , Xiaohan Fei , Ashwin Swaminathan , CJ Taylor , Zhuowen Tu , Paolo Favaro , Stefano Soatto

Existing neural rendering-based text-to-3D-portrait generation methods typically make use of human geometry prior and diffusion models to obtain guidance. However, relying solely on geometry information introduces issues such as the Janus…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Yiqian Wu , Hao Xu , Xiangjun Tang , Xien Chen , Siyu Tang , Zhebin Zhang , Chen Li , Xiaogang Jin

To ease the difficulty of acquiring annotation labels in 3D data, a common method is using unsupervised and open-vocabulary semantic segmentation, which leverage 2D CLIP semantic knowledge. In this paper, unlike previous research that…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Fuyang Yu , Runze Tian , Zhen Wang , Xiaochuan Wang , Xiaohui Liang

Diffusion models excel at capturing complex data distributions, such as those of natural images and proteins. While diffusion models are trained to represent the distribution in the training dataset, we often are more concerned with other…

Detecting 3D objects from multi-view images is a fundamental problem in 3D computer vision. Recently, significant breakthrough has been made in multi-view 3D detection tasks. However, the unprecedented detection performance of these vision…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Linfeng Zhang , Yukang Shi , Hung-Shuo Tai , Zhipeng Zhang , Yuan He , Ke Wang , Kaisheng Ma

With the remarkable advent of text-to-image diffusion models, image editing methods have become more diverse and continue to evolve. A promising recent approach in this realm is Delta Denoising Score (DDS) - an image editing technique based…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Hyelin Nam , Gihyun Kwon , Geon Yeong Park , Jong Chul Ye

We present Piva (Preserving Identity with Variational Score Distillation), a novel optimization-based method for editing images and 3D models based on diffusion models. Specifically, our approach is inspired by the recently proposed method…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Duong H. Le , Tuan Pham , Aniruddha Kembhavi , Stephan Mandt , Wei-Chiu Ma , Jiasen Lu

Text-to-3D, known for its efficient generation methods and expansive creative potential, has garnered significant attention in the AIGC domain. However, the pixel-wise rendering of NeRF and its ray marching light sampling constrain the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Xinhai Li , Huaibin Wang , Kuo-Kun Tseng

Traditional novel view synthesis methods heavily rely on external camera pose estimation tools such as COLMAP, which often introduce computational bottlenecks and propagate errors. To address these challenges, we propose a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Xianben Yang , Yuxuan Li , Tao Wang , Tao Wang , Yi Jin , Yidong Li , Haibin Ling

Recent approaches have shown promises distilling diffusion models into efficient one-step generators. Among them, Distribution Matching Distillation (DMD) produces one-step generators that match their teacher in distribution, without…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Tianwei Yin , Michaël Gharbi , Taesung Park , Richard Zhang , Eli Shechtman , Fredo Durand , William T. Freeman

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Yiwen Chen , Chi Zhang , Xiaofeng Yang , Zhongang Cai , Gang Yu , Lei Yang , Guosheng Lin

Disentangling content and style from a single image, known as content-style decomposition (CSD), enables recontextualization of extracted content and stylization of extracted styles, offering greater creative flexibility in visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Quang-Binh Nguyen , Minh Luu , Quang Nguyen , Anh Tran , Khoi Nguyen

We tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods…

Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Youjia Zhang , Zikai Song , Junqing Yu , Yawei Luo , Wei Yang

Text-to-3D generation has achieved remarkable success via large-scale text-to-image diffusion models. Nevertheless, there is no paradigm for scaling up the methodology to urban scale. Urban scenes, characterized by numerous elements,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Fan Lu , Kwan-Yee Lin , Yan Xu , Hongsheng Li , Guang Chen , Changjun Jiang