English
Related papers

Related papers: From Image Captioning to Visual Storytelling

200 papers

Story Visualization aims to generate images aligned with story prompts, reflecting the coherence of storybooks through visual consistency among characters and scenes.Whereas current approaches exclusively concentrate on characters and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Sitong Su , Litao Guo , Lianli Gao , Heng Tao Shen , Jingkuan Song

This paper presents ViTOC (Vision Transformer and Object-aware Captioner), a novel vision-language model for image captioning that addresses the challenges of accuracy and diversity in generated descriptions. Unlike conventional approaches,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Feiyang Huang

Image captioning is a longstanding problem in the field of computer vision and natural language processing. To date, researchers have produced impressive state-of-the-art performance in the age of deep learning. Most of these…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Zihang Meng , David Yang , Xuefei Cao , Ashish Shah , Ser-Nam Lim

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in…

Computation and Language · Computer Science 2025-06-11 Mohamed Gado , Towhid Taliee , Muhammad Memon , Dmitry Ignatov , Radu Timofte

Story visualization aims to generate a series of images that match the story described in texts, and it requires the generated images to satisfy high quality, alignment with the text description, and consistency in character identities.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Wen Wang , Canyu Zhao , Hao Chen , Zhekai Chen , Kecheng Zheng , Chunhua Shen

We propose a new task, called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-19 Yitong Li , Zhe Gan , Yelong Shen , Jingjing Liu , Yu Cheng , Yuexin Wu , Lawrence Carin , David Carlson , Jianfeng Gao

Online writers and journalism media are increasingly combining visualization (and other multimedia content) with narrative text to create narrative visualizations. Often, however, the two elements are presented independently of one another.…

Human-Computer Interaction · Computer Science 2018-01-09 Ronald Metoyer , Qiyu Zhi , Bart Janczuk , Walter Scheirer

Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their correspondence to the world.…

Computation and Language · Computer Science 2017-10-03 Stephanie Zhou , Alane Suhr , Yoav Artzi

Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human perception process, where the scene is effectively understood…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Kashu Yamazaki , Khoa Vo , Sang Truong , Bhiksha Raj , Ngan Le

Creating meaningful visual narratives through human-AI collaboration requires understanding how text-image intertextuality emerges when textual intentions meet AI-generated visuals. We conducted a three-phase qualitative study with 15…

Human-Computer Interaction · Computer Science 2025-11-06 Mengyao Guo , Kexin Nie , Ze Gao , Black Sun , Xueyang Wang , Jinda Han , Xingting Wu

We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated news articles, image…

Computer Vision and Pattern Recognition · Computer Science 2021-09-15 Fuxiao Liu , Yinghan Wang , Tianlu Wang , Vicente Ordonez

Automatically generating descriptive captions for images is a well-researched area in computer vision. However, existing evaluation approaches focus on measuring the similarity between two sentences disregarding fine-grained semantics of…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Philipp Harzig , Dan Zecha , Rainer Lienhart , Carolin Kaiser , René Schallner

People say, "A picture is worth a thousand words". Then how can we get the rich information out of the image? We argue that by using visual clues to bridge large pretrained vision foundation models and language models, we can do so without…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Yujia Xie , Luowei Zhou , Xiyang Dai , Lu Yuan , Nguyen Bach , Ce Liu , Michael Zeng

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Visual storytelling is the task of generating stories based on a sequence of images. Inspired by the recent works in neural generation focusing on controlling the form of text, this paper explores the idea of generating these stories in…

Computation and Language · Computer Science 2019-06-18 Shrimai Prabhumoye , Khyathi Raghavi Chandu , Ruslan Salakhutdinov , Alan W Black

Visual question answering (VQA) and image captioning require a shared body of general knowledge connecting language and vision. We present a novel approach to improve VQA performance that exploits this connection by jointly generating…

Computer Vision and Pattern Recognition · Computer Science 2020-01-07 Jialin Wu , Zeyuan Hu , Raymond J. Mooney

In the era of evolving artificial intelligence, machines are increasingly emulating human-like capabilities, including visual perception and linguistic expression. Image captioning stands at the intersection of these domains, enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Hrishikesh Singh , Aarti Sharma , Millie Pant

Writing a coherent and engaging story is not easy. Creative writers use their knowledge and worldview to put disjointed elements together to form a coherent storyline, and work and rework iteratively toward perfection. Automated visual…

Computation and Language · Computer Science 2021-07-08 Chi-Yang Hsu , Yun-Wei Chu , Ting-Hao 'Kenneth' Huang , Lun-Wei Ku

Image captioning models require the high-level generalization ability to describe the contents of various images in words. Most existing approaches treat the image-caption pairs equally in their training without considering the differences…

Computer Vision and Pattern Recognition · Computer Science 2022-12-15 Hongkuan Zhang , Saku Sugawara , Akiko Aizawa , Lei Zhou , Ryohei Sasano , Koichi Takeda

Visual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes. Existing autoregressive methods, which rely on previous frame-sentence pairs, struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Sixiao Zheng , Yanwei Fu