中文
相关论文

相关论文: TaleCrafter: Interactive Story Visualization with …

200 篇论文

We present a system to convert any set of images (e.g., a video clip or a photo album) into a storyboard. We aim to create multiple pleasing graphic representations of the content at interactive rates, so the user can explore and find the…

Recent advancements in text-to-image generative models have improved narrative consistency in story visualization. However, current story visualization models often overlook cultural dimensions, resulting in visuals that lack authenticity…

多媒体 · 计算机科学 2025-12-01 Janak Kapuriya , Ali Hatami , Paul Buitelaar

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story with a global consistency across dynamic scenes and characters. Current works still struggle with output images' quality and…

计算机视觉与模式识别 · 计算机科学 2022-09-23 Bowen Li , Thomas Lukasiewicz

The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Chong Mou , Xintao Wang , Liangbin Xie , Yanze Wu , Jian Zhang , Zhongang Qi , Ying Shan , Xiaohu Qie

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

Recent text-to-image (T2I) generative models allow for high-quality synthesis following either text instructions or visual examples. Despite their capabilities, these models face limitations in creating new, detailed creatures within…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Kam Woh Ng , Xiatian Zhu , Yi-Zhe Song , Tao Xiang

Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecified prompts, leading to suboptimal image-text alignment, aesthetics, and quality. We propose a…

计算与语言 · 计算机科学 2025-10-16 Ruibo Chen , Jiacheng Pan , Heng Huang , Zhenheng Yang

Automated plot generation for games enhances the player's experience by providing rich and immersive narrative experience that adapts to the player's actions. Traditional approaches adopt a symbolic narrative planning method which limits…

人机交互 · 计算机科学 2024-11-05 Yi Wang , Qian Zhou , David Ledo

As the text-to-image (T2I) domain progresses, generating text that seamlessly integrates with visual content has garnered significant attention. However, even with accurate text generation, the inability to control font and color can…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Yuxiang Tuo , Yifeng Geng , Liefeng Bo

Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner.…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Ming Tao , Bing-Kun Bao , Hao Tang , Yaowei Wang , Changsheng Xu

Text-to-image (T2I) models have substantially improved image fidelity and prompt adherence, yet their creativity remains constrained by reliance on discrete natural language prompts. When presented with fuzzy prompts such as ``a creative…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Ruixiao Shi , Fu Feng , Yucheng Xie , Xu Yang , Jing Wang , Xin Geng

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

Writing a coherent and engaging story is not easy. Creative writers use their knowledge and worldview to put disjointed elements together to form a coherent storyline, and work and rework iteratively toward perfection. Automated visual…

计算与语言 · 计算机科学 2021-07-08 Chi-Yang Hsu , Yun-Wei Chu , Ting-Hao 'Kenneth' Huang , Lun-Wei Ku

The rapid advancement of text-to-video (T2V) models has revolutionized content creation, yet their commercial potential remains largely untapped. We introduce, for the first time, the task of seamless brand integration in T2V: automatically…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Zihao Zhu , Ruotong Wang , Siwei Lyu , Min Zhang , Baoyuan Wu

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ziyang Liu , Kevin Valencia , Justin Cui

The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts,…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Yiping Wang , Xuehai He , Kuan Wang , Luyao Ma , Jianwei Yang , Shuohang Wang , Simon Shaolei Du , Yelong Shen

Benefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Chun-Mei Feng , Kai Yu , Xinxing Xu , Salman Khan , Rick Siow Mong Goh , Wangmeng Zuo , Yong Liu

Text-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation. Existing methods often partially utilize image inputs, focusing on specific elements like…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Teng-Fang Hsiao , Bo-Kai Ruan , Yi-Lun Wu , Tzu-Ling Lin , Hong-Han Shuai

Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions, enabling vast applications in content generation. While recent advancements have introduced…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Jiun Tian Hoe , Xudong Jiang , Chee Seng Chan , Yap-Peng Tan , Weipeng Hu

Story visualization (SV) is a challenging text-to-image generation task for the difficulty of not only rendering visual details from the text descriptions but also encoding a long-term context across multiple sentences. While prior efforts…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Daechul Ahn , Daneul Kim , Gwangmo Song , Seung Hwan Kim , Honglak Lee , Dongyeop Kang , Jonghyun Choi