中文
相关论文

相关论文: Proportion and Perspective Control for Flow-Based …

200 篇论文

We propose a novel training-free image generation algorithm that precisely controls the occlusion relationships between objects in an image. Existing image generation methods typically rely on prompts to influence occlusion, which often…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xiaohang Zhan , Dingming Liu

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Abdelrahman Eldesokey , Peter Wonka

Text-driven person image generation is an emerging and challenging task in cross-modality image generation. Controllable person image generation promotes a wide range of applications such as digital human interaction and virtual try-on.…

计算机视觉与模式识别 · 计算机科学 2022-11-14 Kaiduo Zhang , Muyi Sun , Jianxin Sun , Binghao Zhao , Kunbo Zhang , Zhenan Sun , Tieniu Tan

We provide a two-way integration for the widely adopted ControlNet by integrating external condition generation algorithms into a single dense prediction method and incorporating its individually trained image generation processes into a…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Yilin Wang , Haiyang Xu , Xiang Zhang , Zeyuan Chen , Zhizhou Sha , Zirui Wang , Zhuowen Tu

Diffusion models have enabled high-quality, conditional image editing capabilities. We propose to expand their arsenal, and demonstrate that off-the-shelf diffusion models can be used for a wide range of cross-domain compositing tasks.…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Roy Hachnochi , Mingrui Zhao , Nadav Orzech , Rinon Gal , Ali Mahdavi-Amiri , Daniel Cohen-Or , Amit Haim Bermano

We introduce Control4D, an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing, notably the inefficiencies of existing 4D representations and the…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Ruizhi Shao , Jingxiang Sun , Cheng Peng , Zerong Zheng , Boyao Zhou , Hongwen Zhang , Yebin Liu

Text-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusion-based editing, where a source image is edited according to a target prompt,…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Xuan Ju , Ailing Zeng , Yuxuan Bian , Shaoteng Liu , Qiang Xu

For an artist or a graphic designer, the spatial layout of a scene is a critical design choice. However, existing text-to-image diffusion models provide limited support for incorporating spatial information. This paper introduces Composite…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Vikram Jamwal , Ramaneswaran S

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Minjung Shin , Hyunin Cho , Sooyeon Go , Jin-Hwa Kim , Youngjung Uh

We propose a novel image editing technique that enables 3D manipulations on single images, such as object rotation and translation. Existing 3D-aware image editing approaches typically rely on synthetic multi-view datasets for training…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Ruicheng Wang , Jianfeng Xiang , Jiaolong Yang , Xin Tong

ControlNets are widely used for adding spatial control to text-to-image diffusion models with different conditions, such as depth maps, scribbles/sketches, and human poses. However, when it comes to controllable video generation,…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Han Lin , Jaemin Cho , Abhay Zala , Mohit Bansal

Controllable image generation has always been one of the core demands in image generation, aiming to create images that are both creative and logical while satisfying additional specified conditions. In the post-AIGC era, controllable…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Guandong Li

Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While recent advances in image generation have enabled high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Kazuaki Mishima , Antoni Bigata Casademunt , Stavros Petridis , Maja Pantic , Kenji Suzuki

Generating realistic animated videos from static images is an important area of research in computer vision. Methods based on physical simulation and motion prediction have achieved notable advances, but they are often limited to specific…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Qiang Wang , Minghua Liu , Junjun Hu , Fan Jiang , Mu Xu

We introduce a novel sketch-to-image tool that aligns with the iterative refinement process of artists. Our tool lets users sketch blocking strokes to coarsely represent the placement and form of objects and detail strokes to refine their…

图形学 · 计算机科学 2024-10-28 Vishnu Sarukkai , Lu Yuan , Mia Tang , Maneesh Agrawala , Kayvon Fatahalian

Controllable image synthesis with user scribbles has gained huge public interest with the recent advent of text-conditioned latent diffusion models. The user scribbles control the color composition while the text prompt provides control…

计算机视觉与模式识别 · 计算机科学 2022-12-01 Jaskirat Singh , Stephen Gould , Liang Zheng

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Diffusion-based image synthesis has attracted extensive attention recently. In particular, ControlNet that uses image-based prompts exhibits powerful capability in image tasks such as canny edge detection and generates images well aligned…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Junjie Yang , Jinze Zhao , Peihao Wang , Zhangyang Wang , Yingbin Liang

Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Changgu Chen , Junwei Shu , Gaoqi He , Changbo Wang , Yang Li

Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score distillation to leverage the 2D diffusion priors to supervise…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Chaohui Yu , Qiang Zhou , Jingliang Li , Zhe Zhang , Zhibin Wang , Fan Wang