中文
相关论文

相关论文: STAGE: Storyboard-Anchored Generation for Cinemati…

200 篇论文

Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and controllability beyond text prompts. To tackle these…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Qinghe Wang , Xiaoyu Shi , Baolu Li , Weikang Bian , Quande Liu , Huchuan Lu , Xintao Wang , Pengfei Wan , Kun Gai , Xu Jia

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yawen Luo , Xiaoyu Shi , Junhao Zhuang , Yutian Chen , Quande Liu , Xintao Wang , Pengfei Wan , Tianfan Xue

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li

We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Quynh Phung , Long Mai , Fabian David Caba Heilbron , Feng Liu , Jia-Bin Huang , Cusuh Ham

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Kaiwen Zhang , Liming Jiang , Angtian Wang , Jacob Zhiyuan Fang , Tiancheng Zhi , Qing Yan , Hao Kang , Xin Lu , Xingang Pan

The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantly rely on sparse conditioning signals such as text prompts…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Zhida Zhang , Jie Ma , Zhan Peng , Haoxue Wu , Yang Han , Jun Liang , Jie Cao , Jing Li

Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Jiaxu Zhang , Tianshu Hu , Yuan Zhang , Zenan Li , Linjie Luo , Guosheng Lin , Xin Chen

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) methods struggle to…

Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Chi Zhang , Yuanzhi Liang , Xi Qiu , Fangqiu Yi , Xuelong Li

We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated in the scene and its spatial relation to the scene, the goal is to synthesize an output…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Cong Xie , Che Wang , Yan Zhang , Ruiqi Yu , Han Zou , Zheng Pan , Zhenpeng Zhan

Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter-shot consistency…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Bingliang Li , Zhenhong Sun , Jiaming Bian , Yuehao Wu , Yifu Wang , Hongdong Li , Yatao Bian , Huadong Mo , Daoyi Dong

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Taewon Kang , Ming C. Lin

This paper introduces StoryAnchors, a unified framework for generating high-quality, multi-scene story frames with strong temporal consistency. The framework employs a bidirectional story generator that integrates both past and future…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Bo Wang , Haoyang Huang , Zhiying Lu , Fengyuan Liu , Guoqing Ma , Jianlong Yuan , Yuan Zhang , Nan Duan , Daxin Jiang

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating…

人工智能 · 计算机科学 2026-04-13 Haobo Hu , Qi Mao , Yuanhang Li , Libiao Jin

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

Recent advances in generative models have made it possible to create high-quality, coherent music, with some systems delivering production-level output. Yet, most existing models focus solely on generating music from scratch, limiting their…

Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven interactions. While prior benchmarks target individual subtasks such as question answering or…

计算与语言 · 计算机科学 2026-05-27 Qiuyu Tian , Zequn Liu , Yiding Li , Fengyi Chen , Youyong Kong , Fan Guo , Yuyao Li , Jinjing Shen , Zhijing Xie , Yiyun Luo , Xin Zhang , Yingce Xia

Video transitions aim to synthesize intermediate frames between two clips, but naive approaches such as linear blending introduce artifacts that limit professional use or break temporal coherence. Traditional techniques (cross-fades,…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Mia Kan , Yilin Liu , Niloy Mitra

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jiamin Wang , Yichen Yao , Xiang Feng , Hang Wu , Yaming Wang , Qingqiu Huang , Yuexin Ma , Xinge Zhu
‹ 上一页 1 2 3 10 下一页 ›