English

Piece it Together: Part-Based Concepting with IP-Priors

Computer Vision and Pattern Recognition 2025-03-14 v1

Abstract

Advanced generative models excel at synthesizing images but often rely on text-based conditioning. Visual designers, however, often work beyond language, directly drawing inspiration from existing visual elements. In many cases, these elements represent only fragments of a potential concept-such as an uniquely structured wing, or a specific hairstyle-serving as inspiration for the artist to explore how they can come together creatively into a coherent whole. Recognizing this need, we introduce a generative framework that seamlessly integrates a partial set of user-provided visual components into a coherent composition while simultaneously sampling the missing parts needed to generate a plausible and complete concept. Our approach builds on a strong and underexplored representation space, extracted from IP-Adapter+, on which we train IP-Prior, a lightweight flow-matching model that synthesizes coherent compositions based on domain-specific priors, enabling diverse and context-aware generations. Additionally, we present a LoRA-based fine-tuning strategy that significantly improves prompt adherence in IP-Adapter+ for a given task, addressing its common trade-off between reconstruction quality and prompt adherence.

Keywords

Cite

@article{arxiv.2503.10365,
  title  = {Piece it Together: Part-Based Concepting with IP-Priors},
  author = {Elad Richardson and Kfir Goldberg and Yuval Alaluf and Daniel Cohen-Or},
  journal= {arXiv preprint arXiv:2503.10365},
  year   = {2025}
}

Comments

Project page available at https://eladrich.github.io/PiT/

R2 v1 2026-06-28T22:19:03.353Z