English
Related papers

Related papers: Customized Visual Storytelling with Unified Multim…

200 papers

Preference-conditioned image generation seeks to adapt generative models to individual users, producing outputs that reflect personal aesthetic choices beyond the given textual prompt. Despite recent progress, existing approaches either…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Wenyi Mo , Tianyu Zhang , Yalong Bai , Ligong Han , Ying Ba , Dimitris N. Metaxas

Automated movie creation requires coordinating multiple characters, modalities, and narrative elements across extended sequences -- a challenge that existing end-to-end approaches struggle to address effectively. We present…

Multimedia · Computer Science 2026-04-28 Tianyidan Xie , Zhentao Huang , Mingjie Wang , Xin Huang , Jun Zhou , Minglun Gong , Zili Yi

Storyline visualization has emerged as an innovative method for illustrating the development and changes in stories across various domains. Traditional approaches typically represent stories with one line per character, progressing from…

Human-Computer Interaction · Computer Science 2024-08-06 Haonan Yao , Lixiang Zhao , Boyuan Chen , Kaiwen Li , Hai-Ning Liang , Lingyun Yu

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Qian Liang , Yujia Wu , Kuncheng Li , Jiwei Wei , Shiyuan He , Jinyu Guo , Ning Xie

Generating realistic and diverse trajectories is a critical challenge in autonomous driving simulation. While Large Language Models (LLMs) show promise, existing methods often rely on structured data like vectorized maps, which fail to…

Artificial Intelligence · Computer Science 2026-03-06 Mingxuan Mu , Guo Yang , Lei Chen , Ping Wu , Jianxun Cui

Recognizing characters and predicting speakers of dialogue are critical for comic processing tasks, such as voice generation or translation. However, because characters vary by comic title, supervised learning approaches like training…

Multimedia · Computer Science 2024-09-06 Yingxuan Li , Ryota Hinami , Kiyoharu Aizawa , Yusuke Matsui

Recent generative models have demonstrated impressive capabilities in generating realistic and visually pleasing images grounded on textual prompts. Nevertheless, a significant challenge remains in applying these models for the more…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Xiaoqian Shen , Mohamed Elhoseiny

Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Mingzhe Zheng , Yongqi Xu , Haojian Huang , Xuran Ma , Yexin Liu , Wenjie Shu , Yatian Pang , Feilong Tang , Qifeng Chen , Harry Yang , Ser-Nam Lim

As Vision-Language Models (VLMs) achieve widespread deployment across diverse cultural contexts, ensuring their cultural competence becomes critical for responsible AI systems. While prior work has evaluated cultural awareness in text-only…

Computation and Language · Computer Science 2025-08-26 Arka Mukherjee , Shreya Ghosh

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Taewon Kang , Ming C. Lin

Visual storytelling aims to generate compelling narratives from image sequences. Existing models often focus on enhancing the representation of the image sequence, e.g., with external knowledge sources or advanced graph structures. Despite…

Computation and Language · Computer Science 2023-10-19 Danyang Liu , Mirella Lapata , Frank Keller

Story visualization requires generating sequential imagery that aligns semantically with evolving narratives while maintaining rigorous consistency in character identity and visual style. However, existing methodologies often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Jianzhang Zhang , Yijing Tian , Jiwang Qu , Chuang Liu

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inaccurate dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Ayan Banerjee , Josep Llados , Umapada Pal , Anjan Dutta

Storytelling aims to generate reasonable and vivid narratives based on an ordered image stream. The fidelity to the image story theme and the divergence of story plots attract readers to keep reading. Previous works iteratively improved the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Chuanqi Zang , Jiji Tang , Rongsheng Zhang , Zeng Zhao , Tangjie Lv , Mingtao Pei , Wei Liang

Anomaly detection is vital in various industrial scenarios, including the identification of unusual patterns in production lines and the detection of manufacturing defects for quality control. Existing techniques tend to be specialized in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Xiaohao Xu , Yunkang Cao , Huaxin Zhang , Nong Sang , Xiaonan Huang

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

We propose MLLM-VADStory, a novel domain knowledge-guided multimodal large language models (MLLM) framework to systematically quantify and generate insights for video ad storyline understanding at scale. The framework is centered on the…

Multimedia · Computer Science 2026-01-14 Jasmine Yang , Poppy Zhang , Shawndra Hill

The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantly rely on sparse conditioning signals such as text prompts…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Zhida Zhang , Jie Ma , Zhan Peng , Haoxue Wu , Yang Han , Jun Liang , Jie Cao , Jing Li

Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero-shot approaches, particularly those leveraging diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Yiheng Lin , Shifang Zhao , Ting Liu , Xiaochao Qu , Luoqi Liu , Yao Zhao , Yunchao Wei