English
Related papers

Related papers: DisenStudio: Customized Multi-subject Text-to-Vide…

200 papers

Creating 3D assets that follow the texture and geometry style of existing ones is often desirable or even inevitable in practical applications like video gaming and virtual reality. While impressive progress has been made in generating 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Zefan Qu , Zhenwei Wang , Haoyuan Wang , Ke Xu , Gerhard Hancke , Rynson W. H. Lau

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Bo Peng , Xinyuan Chen , Yaohui Wang , Chaochao Lu , Yu Qiao

Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of guidance. Existing methods often struggle to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Zicheng Duan , Yuxuan Ding , Chenhui Gou , Ziqin Zhou , Ethan Smith , Lingqiao Liu

Multi-subject personalized generation presents unique challenges in maintaining identity fidelity and semantic coherence when synthesizing images conditioned on multiple reference subjects. Existing methods often suffer from identity…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Dong She , Siming Fu , Mushui Liu , Qiaoqiao Jin , Hualiang Wang , Mu Liu , Jidong Jiang

Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Edoardo A. Dominici , Thomas Deixelberger , Konstantinos Vardis , Markus Steinberger

Recently, the multimedia community has witnessed the rise of diffusion models trained on large-scale multi-modal data for visual content creation, particularly in the field of text-to-image generation. In this paper, we propose a new task…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Jingwen Chen , Yingwei Pan , Ting Yao , Tao Mei

Personalizing text-to-image models to generate images of specific subjects across diverse scenes and styles is a rapidly advancing field. Current approaches often face challenges in maintaining a balance between identity preservation and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Or Patashnik , Rinon Gal , Daniil Ostashev , Sergey Tulyakov , Kfir Aberman , Daniel Cohen-Or

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Xiangpeng Yang , Linchao Zhu , Hehe Fan , Yi Yang

Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. There have been a number of recent works dealing with learning and composing multiple objects…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Sonali Godavarthy , Matthias Neuwirth-Trapp , Tim-Felix Faasch , Maarten Bieshaar , Michael Moeller , Danda Pani Paudel

Diffusion-based models have achieved state-of-the-art performance on text-to-image synthesis tasks. However, one critical limitation of these models is the low fidelity of generated images with respect to the text description, such as…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Qiucheng Wu , Yujian Liu , Handong Zhao , Trung Bui , Zhe Lin , Yang Zhang , Shiyu Chang

Recent advances in text-to-image generation with diffusion models present transformative capabilities in image quality. However, user controllability of the generated image, and fast adaptation to new tasks still remains an open challenge,…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Omer Bar-Tal , Lior Yariv , Yaron Lipman , Tali Dekel

We present DreamBooth3D, an approach to personalize text-to-3D generative models from as few as 3-6 casually captured images of a subject. Our approach combines recent advances in personalizing text-to-image models (DreamBooth) with…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Amit Raj , Srinivas Kaza , Ben Poole , Michael Niemeyer , Nataniel Ruiz , Ben Mildenhall , Shiran Zada , Kfir Aberman , Michael Rubinstein , Jonathan Barron , Yuanzhen Li , Varun Jampani

Text-to-image customization, which aims to synthesize text-driven images for the given subjects, has recently revolutionized content creation. Existing works follow the pseudo-word paradigm, i.e., represent the given subjects as…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Mengqi Huang , Zhendong Mao , Mingcong Liu , Qian He , Yongdong Zhang

Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Jingyuan Zhu , Huimin Ma , Jiansheng Chen , Jian Yuan

Given a text and an image of a specific subject, text-to-image customization aims to generate new images that align with both the text and the subject's appearance. Existing works follow the pseudo-word paradigm, which represents the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zhendong Mao , Mengqi Huang , Fei Ding , Mingcong Liu , Qian He , Yongdong Zhang

Text-to-image diffusion models have shown remarkable success in generating personalized subjects based on a few reference images. However, current methods often fail when generating multiple subjects simultaneously, resulting in mixed…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Sangwon Jang , Jaehyeong Jo , Kimin Lee , Sung Ju Hwang

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yuancheng Xu , Wenqi Xian , Li Ma , Julien Philip , Ahmet Levent Taşel , Yiwei Zhao , Ryan Burgert , Mingming He , Oliver Hermann , Oliver Pilarski , Rahul Garg , Paul Debevec , Ning Yu

Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Chenyang Si , Weichen Fan , Zhengyao Lv , Ziqi Huang , Yu Qiao , Ziwei Liu

Leveraging large-scale image-text datasets and advancements in diffusion models, text-driven generative models have made remarkable strides in the field of image generation and editing. This study explores the potential of extending the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Fu-Yun Wang , Wenshuo Chen , Guanglu Song , Han-Jia Ye , Yu Liu , Hongsheng Li
‹ Prev 1 3 4 5 6 7 10 Next ›