English
Related papers

Related papers: CANVAS: Continuity-Aware Narratives via Visual Age…

200 papers

We present Story2Board, a training-free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 David Dinkevich , Matan Levy , Omri Avrahami , Dvir Samuel , Dani Lischinski

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Peixuan Zhang , Zijian Jia , Kaiqi Liu , Shuchen Weng , Si Li , Boxin Shi

Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter-shot consistency…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Bingliang Li , Zhenhong Sun , Jiaming Bian , Yuehao Wu , Yifu Wang , Hongdong Li , Yatao Bian , Huadong Mo , Daoyi Dong

A storyboard is a sequence of images to illustrate a story containing multiple sentences, which has been a key process to create different story products. In this paper, we tackle a new multimedia task of automatic storyboard creation to…

Machine Learning · Computer Science 2019-12-02 Shizhe Chen , Bei Liu , Jianlong Fu , Ruihua Song , Qin Jin , Pingping Lin , Xiaoyu Qi , Chunting Wang , Jin Zhou

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Hailong Yan , Shice Liu , Tao Wang , Xiangtao Zhang , Yijie Zhong , Jinwei Chen , Le Zhang , Bo Li

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating…

Artificial Intelligence · Computer Science 2026-04-13 Haobo Hu , Qi Mao , Yuanhang Li , Libiao Jin

Human writers often begin their stories with an overarching mental scene, where they envision the interactions between characters and their environment. Inspired by this creative process, we propose a novel approach to long-form story…

Computation and Language · Computer Science 2026-03-20 Zehao Chen , Rong Pan , Haoran Li

Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Jaskirat Singh , Junshen Kevin Chen , Jonas Kohler , Michael Cohen

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story with a global consistency across dynamic scenes and characters. Current works still struggle with output images' quality and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-23 Bowen Li , Thomas Lukasiewicz

With the maturity of visual detection techniques, we are more ambitious in describing visual content with open-vocabulary, fine-grained and free-form language, i.e., the task of image captioning. In particular, we are interested in…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zheng-Jun Zha , Daqing Liu , Hanwang Zhang , Yongdong Zhang , Feng Wu

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li

Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Chi Zhang , Yuanzhi Liang , Xi Qiu , Fangqiu Yi , Xuelong Li

Story visualization is an under-explored task that falls at the intersection of many important research directions in both computer vision and natural language processing. In this task, given a series of natural language captions which…

Computation and Language · Computer Science 2021-05-24 Adyasha Maharana , Darryl Hannan , Mohit Bansal

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Xinlei Yin , Xiulian Peng , Xiao Li , Zhiwei Xiong , Yan Lu

While diffusion models generate high-fidelity video clips, transforming them into coherent storytelling engines remains challenging. Current agentic pipelines automate this via chained modules but suffer from semantic drift and cascading…

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Kaiwen Zhang , Liming Jiang , Angtian Wang , Jacob Zhiyuan Fang , Tiancheng Zhi , Qing Yan , Hao Kang , Xin Lu , Xingang Pan

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Zhongyu Yang , Zuhao Yang , Shuo Zhan , Tan Yue , Wei Pang , Yingfang Yuan

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence…

Computation and Language · Computer Science 2025-08-21 Admitos Passadakis , Yingjin Song , Albert Gatt

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar