中文
相关论文

相关论文: ComposeAnything: Composite Object Priors for Text-…

200 篇论文

Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Jisheng Chu , Wenrui Li , Rui Zhao , Wangmeng Zuo , Shifeng Chen , Xiaopeng Fan

Text-to-image (T2I) models enable rapid concept design, making them widely used in AI-driven design. While recent studies focus on generating semantic and stylistic variations of given design concepts, functional coherence--the integration…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Hyeonjeong Ha , Xiaomeng Jin , Jeonghwan Kim , Jiateng Liu , Zhenhailong Wang , Khanh Duy Nguyen , Ansel Blume , Nanyun Peng , Kai-Wei Chang , Heng Ji

Text-to-image diffusion-based generative models have the stunning ability to generate photo-realistic images and achieve state-of-the-art low FID scores on challenging image generation benchmarks. However, one of the primary failure modes…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Arman Zarei , Keivan Rezaei , Samyadeep Basu , Mehrdad Saberi , Mazda Moayeri , Priyatham Kattakinda , Soheil Feizi

In this paper, we tackle a new task of 3D object synthesis, where a 3D model is composited with another object category to create a novel 3D model. However, most existing text/image/3D-to-3D methods struggle to effectively integrate…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Zeren Xiong , Zikun Chen , Zedong Zhang , Xiang Li , Ying Tai , Jian Yang , Jun Li

We present Material Anything, a fully-automated, unified diffusion framework designed to generate physically-based materials for 3D objects. Unlike existing methods that rely on complex pipelines or case-specific optimizations, Material…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Xin Huang , Tengfei Wang , Ziwei Liu , Qing Wang

Recent years have witnessed the substantial progress of large-scale models across various domains, such as natural language processing and computer vision, facilitating the expression of concrete concepts. Unlike concrete concepts that are…

计算机视觉与模式识别 · 计算机科学 2023-09-28 Jiayi Liao , Xu Chen , Qiang Fu , Lun Du , Xiangnan He , Xiang Wang , Shi Han , Dongmei Zhang

Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints-notably generative numeracy, attribute binding, and part-level relations. To address…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Yu Huo , Siyu Zhang , Kun Zeng , Haoyue Liu , Owen Lee , Junlin Chen , Yuquan Lu , Yifu Guo , Yaodong Liang , Xiaoying Tang

In this paper, we study Text-to-3D content generation leveraging 2D diffusion priors to enhance the quality and detail of the generated 3D models. Recent progress (Magic3D) in text-to-3D has shown that employing high-resolution (e.g., 512 x…

计算机视觉与模式识别 · 计算机科学 2023-08-01 Jinbo Wu , Xiaobo Gao , Xing Liu , Zhengyang Shen , Chen Zhao , Haocheng Feng , Jingtuo Liu , Errui Ding

Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent generative image…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Gemma Canet Tarrés , Zhe Lin , Zhifei Zhang , Jianming Zhang , Yizhi Song , Dan Ruta , Andrew Gilbert , John Collomosse , Soo Ye Kim

Despite significant progress, controlled generation of complex images with interacting people remains difficult. Existing layout generation methods fall short of synthesizing realistic person instances; while pose-guided generation…

计算机视觉与模式识别 · 计算机科学 2020-08-31 Weidong Yin , Ziwei Liu , Leonid Sigal

Despite the ability of text-to-image models to generate high-quality, realistic, and diverse images, they face challenges in compositional generation, often struggling to accurately represent details specified in the input prompt. A…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Parham Rezaei , Arash Marioriyad , Mahdieh Soleymani Baghshah , Mohammad Hossein Rohban

Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quantities, attributes, and visual features. Preventing leakage…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Omer Dahary , Yehonathan Cohen , Or Patashnik , Kfir Aberman , Daniel Cohen-Or

Layer compositing is one of the most popular image editing workflows among both amateurs and professionals. Motivated by the success of diffusion models, we explore layer compositing from a layered image generation perspective. Instead of…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Xinyang Zhang , Wentian Zhao , Xin Lu , Jeff Chien

Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability. This work offers a new generation paradigm that allows flexible control of the output image,…

计算机视觉与模式识别 · 计算机科学 2023-02-23 Lianghua Huang , Di Chen , Yu Liu , Yujun Shen , Deli Zhao , Jingren Zhou

Shape primitive abstraction, which decomposes complex 3D shapes into simple geometric elements, plays a crucial role in human visual cognition and has broad applications in computer vision and graphics. While recent advances in 3D content…

图形学 · 计算机科学 2025-05-08 Jingwen Ye , Yuze He , Yanning Zhou , Yiqin Zhu , Kaiwen Xiao , Yong-Jin Liu , Wei Yang , Xiao Han

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that grasping the cross-modality alignment is not a necessity for a…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Shuailei Ma , Kecheng Zheng , Ying Wei , Wei Wu , Fan Lu , Yifei Zhang , Chen-Wei Xie , Biao Gong , Jiapeng Zhu , Yujun Shen

Contrastively trained vision-language models have achieved remarkable progress in vision and language representation learning, leading to state-of-the-art models for various downstream multimodal tasks. However, recent research has…

计算与语言 · 计算机科学 2023-10-26 Harman Singh , Pengchuan Zhang , Qifan Wang , Mengjiao Wang , Wenhan Xiong , Jingfei Du , Yu Chen

Text-to-image generation is a significant domain in modern computer vision and has achieved substantial improvements through the evolution of generative architectures. Among these, there are diffusion-based models that have demonstrated…

Generating high-quality 3D assets from a given image is highly desirable in various applications such as AR/VR. Recent advances in single-image 3D generation explore feed-forward models that learn to infer the 3D model of an object without…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yongwei Chen , Tengfei Wang , Tong Wu , Xingang Pan , Kui Jia , Ziwei Liu

Despite remarkable progress in Text-to-Image models, many real-world applications require generating coherent image sets with diverse consistency requirements. Existing consistent methods often focus on a specific domain with specific…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Chengyou Jia , Xin Shen , Zhuohang Dang , Zhuohang Dang , Changliang Xia , Weijia Wu , Xinyu Zhang , Hangwei Qian , Ivor W. Tsang , Minnan Luo