中文
相关论文

相关论文: Storybooth: Training-free Multi-Subject Consistenc…

200 篇论文

Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Jingyuan Zhu , Huimin Ma , Jiansheng Chen , Jian Yuan

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Yuming Jiang , Tianxing Wu , Shuai Yang , Chenyang Si , Dahua Lin , Yu Qiao , Chen Change Loy , Ziwei Liu

In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have explored training-free methods to enhance the consistency of…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Mengyu Wang , Henghui Ding , Jianing Peng , Yao Zhao , Yunpeng Chen , Yunchao Wei

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Hong Chen , Xin Wang , Yipeng Zhang , Yuwei Zhou , Zeyang Zhang , Siao Tang , Wenwu Zhu

This paper introduces StoryAnchors, a unified framework for generating high-quality, multi-scene story frames with strong temporal consistency. The framework employs a bidirectional story generator that integrates both past and future…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Bo Wang , Haoyang Huang , Zhiying Lu , Fengyuan Liu , Guoqing Ma , Jianlong Yuan , Yuan Zhang , Nan Duan , Daxin Jiang

Personalizing text-to-image models using a limited set of images for a specific object has been explored in subject-specific image generation. However, existing methods often face challenges in aligning with text prompts due to overfitting…

计算机视觉与模式识别 · 计算机科学 2024-02-16 Daewon Chae , Nokyung Park , Jinkyu Kim , Kimin Lee

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generation begins to attract the attention of related research…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Junhao Cheng , Xi Lu , Hanhui Li , Khun Loun Zai , Baiqiao Yin , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

While modern diffusion models excel at generating diverse single images, extending this to sequential generation reveals a fundamental challenge: balancing narrative dynamism with multi-character coherence. Existing methods often falter at…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Qi Zhao , Jun Chen , Ivor Tsang , Guang Dai

Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the…

计算与语言 · 计算机科学 2020-11-02 Ruize Wang , Zhongyu Wei , Ying Cheng , Piji Li , Haijun Shan , Ji Zhang , Qi Zhang , Xuanjing Huang

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Tao Liu , Kai Wang , Senmao Li , Joost van de Weijer , Fahad Shahbaz Khan , Shiqi Yang , Yaxing Wang , Jian Yang , Ming-Ming Cheng

The audiovisual industry is undergoing a profound transformation as it is integrating AI developments not only to automate routine tasks but also to inspire new forms of art. This paper addresses the problem of producing a virtually…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Ruben Pascual , Mikel Sesma-Sara , Aranzazu Jurio , Daniel Paternain , Mikel Galar

Personalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Qihan Huang , Siming Fu , Jinlong Liu , Hao Jiang , Yipeng Yu , Jie Song

Recent advancements in text-to-image diffusion models have demonstrated their remarkable capability to generate high-quality images from textual prompts. However, increasing research indicates that these models memorize and replicate images…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Jie Ren , Yaxin Li , Shenglai Zeng , Han Xu , Lingjuan Lyu , Yue Xing , Jiliang Tang

Diffusion models excel at text-to-image generation, especially in subject-driven generation for personalized images. However, existing methods are inefficient due to the subject-specific fine-tuning, which is computationally intensive and…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Guangxuan Xiao , Tianwei Yin , William T. Freeman , Frédo Durand , Song Han

Leveraging the generative ability of image diffusion models offers great potential for zero-shot video-to-video translation. The key lies in how to maintain temporal consistency across generated video frames by image diffusion models.…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Yuxiang Bao , Di Qiu , Guoliang Kang , Baochang Zhang , Bo Jin , Kaiye Wang , Pengfei Yan

In text-to-image diffusion models, the cross-attention map of each text token indicates the specific image regions attended. Comparing these maps of syntactically related tokens provides insights into how well the generated image reflects…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Jeeyung Kim , Erfan Esmaeili , Qiang Qiu

Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Kiymet Akdemir , Tahira Kazimi , Pinar Yanardag

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

This paper introduces MultiBooth, a novel and efficient technique for multi-concept customization in image generation from text. Despite the significant advancements in customized generation methods, particularly with the success of…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Chenyang Zhu , Kai Li , Yue Ma , Chunming He , Xiu Li