English
Related papers

Related papers: AnyMS: Bottom-up Attention Decoupling for Layout-g…

200 papers

Multi-ID customization is an interesting topic in computer vision and attracts considerable attention recently. Given the ID images of multiple individuals, its purpose is to generate a customized image that seamlessly integrates them while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Jiawei Lin , Guanlong Jiao , Jianjin Xu

Recent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Xierui Wang , Siming Fu , Qihan Huang , Wanggui He , Hao Jiang

Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Fei Peng , Junqiang Wu , Yan Li , Tingting Gao , Di Zhang , Huiyuan Fu

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Pengxiang Cai , Mengyang Li

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Junjie Hu , Tianyang Han , Kai Ma , Jialin Gao , Song Yang , Xianhua He , Junfeng Luo , Xiaoming Wei , Wenqiang Zhang

Current multi-subject customization approaches encounter two critical challenges: the difficulty in acquiring diverse multi-subject training data, and attribute entanglement across different subjects. To bridge these gaps, we propose MUSAR…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Zinan Guo , Pengze Zhang , Yanze Wu , Chong Mou , Songtao Zhao , Qian He

Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Omer Dahary , Or Patashnik , Kfir Aberman , Daniel Cohen-Or

Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Lincan Cai , Jingxuan Kang , Shuang Li , Wenxuan Ma , Binhui Xie , Zhida Qin , Jian Liang

Multi-view learning primarily aims to fuse multiple features to describe data comprehensively. Most prior studies implicitly assume that different views share similar dimensions. In practice, however, severe dimensional disparities often…

Machine Learning · Computer Science 2026-04-01 Cai Xu , Changhao Sun , Ziyu Guan , Wei Zhao

Diffusion models excel at text-to-image generation, especially in subject-driven generation for personalized images. However, existing methods are inefficient due to the subject-specific fine-tuning, which is computationally intensive and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Guangxuan Xiao , Tianwei Yin , William T. Freeman , Frédo Durand , Song Han

Despite the progress of image segmentation for accurate visual entity segmentation, completing the diverse requirements of image editing applications for different-level region-of-interest selections remains unsolved. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Lu Qi , Jason Kuen , Weidong Guo , Jiuxiang Gu , Zhe Lin , Bo Du , Yu Xu , Ming-Hsuan Yang

Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yu Xu , Fan Tang , You Wu , Lin Gao , Oliver Deussen , Hongbin Yan , Jintao Li , Juan Cao , Tong-Yee Lee

The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Han Yang , Chuanguang Yang , Qiuli Wang , Zhulin An , Weilun Feng , Libo Huang , Yongjun Xu

Text-to-image (T2I) customization empowers users to adapt the T2I diffusion model to new concepts absent in the pre-training dataset. On this basis, capturing multiple new concepts from a single image has emerged as a new task, allowing the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Junjie Shentu , Matthew Watson , Noura Al Moubayed

Synthesizing images with user-specified subjects has received growing attention due to its practical applications. Despite the recent success in single subject customization, existing algorithms suffer from high training cost and low…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Zhiheng Liu , Yifei Zhang , Yujun Shen , Kecheng Zheng , Kai Zhu , Ruili Feng , Yu Liu , Deli Zhao , Jingren Zhou , Yang Cao

Multi-subject personalized generation presents unique challenges in maintaining identity fidelity and semantic coherence when synthesizing images conditioned on multiple reference subjects. Existing methods often suffer from identity…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Dong She , Siming Fu , Mushui Liu , Qiaoqiao Jin , Hualiang Wang , Mu Liu , Jidong Jiang

Recent advances in text-to-image diffusion models spurred research on personalization, i.e., a customized image synthesis, of subjects within reference images. Although existing personalization methods are able to alter the subjects'…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yongjin Choi , Chanhun Park , Seung Jun Baek

Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance alignment efficiency, primarily through minimal-data…

Computation and Language · Computer Science 2025-06-19 Hao Chen , Haoze Li , Zhiqing Xiao , Lirong Gao , Qi Zhang , Xiaomeng Hu , Ningtao Wang , Xing Fu , Junbo Zhao

Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quantities, attributes, and visual features. Preventing leakage…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Omer Dahary , Yehonathan Cohen , Or Patashnik , Kfir Aberman , Daniel Cohen-Or

Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Context-Learning based…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Tao Wu , Yibo Jiang , Yehao Lu , Zhizhong Wang , Zeyi Huang , Zequn Qin , Xi Li
‹ Prev 1 2 3 10 Next ›