English
Related papers

Related papers: Proportion and Perspective Control for Flow-Based …

200 papers

Text-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face certain limitations when dealing with personalized image…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Yuhong Zhang , Han Wang , Yiwen Wang , Rong Xie , Li Song

In recent years, significant progress has been made in the development of text-to-image generation models. However, these models still face limitations when it comes to achieving full controllability during the generation process. Often,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Salaheldin Mohamed

While perspective is a well-studied topic in art, it is generally taken for granted in images. However, for the recent wave of high-quality image synthesis methods such as latent diffusion models, perspective accuracy is not an explicit…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Rishi Upadhyay , Howard Zhang , Yunhao Ba , Ethan Yang , Blake Gella , Sicheng Jiang , Alex Wong , Achuta Kadambi

Diffusion models have recently become the de-facto approach for generative modeling in the 2D domain. However, extending diffusion models to 3D is challenging due to the difficulties in acquiring 3D ground truth data for training. On the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Jiatao Gu , Qingzhe Gao , Shuangfei Zhai , Baoquan Chen , Lingjie Liu , Josh Susskind

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Can Qin , Shu Zhang , Ning Yu , Yihao Feng , Xinyi Yang , Yingbo Zhou , Huan Wang , Juan Carlos Niebles , Caiming Xiong , Silvio Savarese , Stefano Ermon , Yun Fu , Ran Xu

Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Ruichen Wang , Zekang Chen , Chen Chen , Jian Ma , Haonan Lu , Xiaodong Lin

The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Andrey Palaev , Adil Khan , Syed M. Ahsan Kazmi

Large-scale generative models are capable of producing high-quality images from detailed text descriptions. However, many aspects of an image are difficult or impossible to convey through text. We introduce self-guidance, a method that…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Dave Epstein , Allan Jabri , Ben Poole , Alexei A. Efros , Aleksander Holynski

Diffusion models offer unprecedented image generation power given just a text prompt. While emerging approaches for controlling diffusion models have enabled users to specify the desired spatial layouts of the generated content, they cannot…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Yunxiang Zhang , Nan Wu , Connor Z. Lin , Gordon Wetzstein , Qi Sun

Recent approaches such as ControlNet offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However, auxiliary modules have to be trained for each type of spatial condition, model architecture, and checkpoint,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Sicheng Mo , Fangzhou Mu , Kuan Heng Lin , Yanli Liu , Bochen Guan , Yin Li , Bolei Zhou

Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Omri Avrahami , Thomas Hayes , Oran Gafni , Sonal Gupta , Yaniv Taigman , Devi Parikh , Dani Lischinski , Ohad Fried , Xi Yin

Recent diffusion-based generative models show promise in their ability to generate text images, but limitations in specifying the styles of the generated texts render them insufficient in the realm of typographic design. This paper proposes…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 KhayTze Peong , Seiichi Uchida , Daichi Haraguchi

We provide an attention-level control method for the task of coupled image generation, where "coupled" means that multiple simultaneously generated images are expected to have the same or very similar backgrounds. While backgrounds coupled,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Chenfei Yuan , Nanshan Jia , Hangqi Li , Peter W. Glynn , Zeyu Zheng

Our ability to sample realistic natural images, particularly faces, has advanced by leaps and bounds in recent years, yet our ability to exert fine-tuned control over the generative process has lagged behind. If this new technology is to…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Marek Kowalski , Stephan J. Garbin , Virginia Estellers , Tadas Baltrušaitis , Matthew Johnson , Jamie Shotton

Manipulation of facial images to meet specific controls such as pose, expression, and lighting, also known as face rigging, is a complex task in computer vision. Existing methods are limited by their reliance on image datasets, which…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Wooseok Jang , Youngjun Hong , Geonho Cha , Seungryong Kim

We introduce MVControl, a novel neural network architecture that enhances existing pre-trained multi-view 2D diffusion models by incorporating additional input conditions, e.g. edge maps. Our approach enables the generation of controllable…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zhiqi Li , Yiming Chen , Lingzhe Zhao , Peidong Liu

Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Zeju Qiu , Weiyang Liu , Haiwen Feng , Yuxuan Xue , Yao Feng , Zhen Liu , Dan Zhang , Adrian Weller , Bernhard Schölkopf

Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Edurne Bernal-Berdun , Ana Serrano , Belen Masia , Matheus Gadelha , Yannick Hold-Geoffroy , Xin Sun , Diego Gutierrez

Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Amadou S. Sangare , Adrien Maglo , Mohamed Chaouch , Bertrand Luvison

ControlNet has enabled detailed spatial control in text-to-image diffusion models by incorporating additional visual conditions such as depth or edge maps. However, its effectiveness heavily depends on the availability of visual conditions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Woosung Joung , Daewon Chae , Jinkyu Kim