English
Related papers

Related papers: Detecting and Grounding Important Characters in Vi…

200 papers

One of the primary challenges of visual storytelling is developing techniques that can maintain the context of the story over long event sequences to generate human-like stories. In this paper, we propose a hierarchical deep learning…

Computer Vision and Pattern Recognition · Computer Science 2019-09-30 Md Sultan Al Nahian , Tasmia Tasrin , Sagar Gandhi , Ryan Gaines , Brent Harrison

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence…

Computation and Language · Computer Science 2025-08-21 Admitos Passadakis , Yingjin Song , Albert Gatt

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Ali Furkan Biten , Ruben Tito , Andres Mafla , Lluis Gomez , Marçal Rusiñol , Ernest Valveny , C. V. Jawahar , Dimosthenis Karatzas

Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through grounding of characters,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Daniel A. P. Oliveira , David Martins de Matos

We introduce the first dataset for human edits of machine-generated visual stories and explore how these collected edits may be used for the visual story post-editing task. The dataset, VIST-Edit, includes 14,905 human edited versions of…

Computation and Language · Computer Science 2019-06-06 Ting-Yao Hsu , Chieh-Yang Huang , Yen-Chia Hsu , Ting-Hao 'Kenneth' Huang

Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Georgios Pantazopoulos , Eda B. Özyiğit

Current image generation models struggle to reliably produce well-formed visual text. In this paper, we investigate a key contributing factor: popular text-to-image models lack character-level input features, making it much harder to…

Computation and Language · Computer Science 2023-05-04 Rosanne Liu , Dan Garrette , Chitwan Saharia , William Chan , Adam Roberts , Sharan Narang , Irina Blok , RJ Mical , Mohammad Norouzi , Noah Constant

Visual dialog is challenging since it needs to answer a series of coherent questions based on understanding the visual environment. How to ground related visual objects is one of the key problems. Previous studies utilize the question and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Feilong Chen , Xiuyi Chen , Can Xu , Daxin Jiang

Narrative inquiry has been one of the prominent application domains for the analysis of human experience, aiming to know more about the complexity of human society. However, researchers are often required to transform various forms of data…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Runtong Wu , Jiayao Song , Fei Teng , Xianhao Ren , Yuyan Gao , Kailun Yang

We define "visual story-writing" as using visual representations of story elements to support writing and revising narrative texts. To demonstrate this approach, we developed a text editor that automatically visualizes a graph of entity…

Human-Computer Interaction · Computer Science 2025-08-01 Damien Masson , Zixin Zhao , Fanny Chevalier

Existing visual explanation generating agents learn to fluently justify a class prediction. However, they may mention visual attributes which reflect a strong class prior, although the evidence may not actually be in the image. This is…

Computer Vision and Pattern Recognition · Computer Science 2018-08-03 Lisa Anne Hendricks , Ronghang Hu , Trevor Darrell , Zeynep Akata

Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable attention to defining…

Computation and Language · Computer Science 2025-04-22 Sabyasachee Baruah , Shrikanth Narayanan

Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the…

Computation and Language · Computer Science 2020-11-02 Ruize Wang , Zhongyu Wei , Ying Cheng , Piji Li , Haijun Shan , Ji Zhang , Qi Zhang , Xuanjing Huang

Visual storytelling aims to generate a narrative based on a sequence of images, necessitating both vision-language alignment and coherent story generation. Most existing solutions predominantly depend on paired image-text training data,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Yuechen Wang , Wengang Zhou , Zhenbo Lu , Houqiang Li

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Claire Yuqing Cui , Apoorv Khandelwal , Yoav Artzi , Noah Snavely , Hadar Averbuch-Elor

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story with a global consistency across dynamic scenes and characters. Current works still struggle with output images' quality and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-23 Bowen Li , Thomas Lukasiewicz

Autonomous agents that can engage in social interactions witha human is the ultimate goal of a myriad of applications. A keychallenge in the design of these applications is to define the socialbehavior of the agent, which requires extensive…

Human-Computer Interaction · Computer Science 2021-07-20 Diogo S. Carvalho , Joana Campos , Manuel Guimarães , Ana Antunes , João Dias , Pedro A. Santos

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-23 Han Hu , Chengquan Zhang , Yuxuan Luo , Yuzhuo Wang , Junyu Han , Errui Ding

Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs…

Human-Computer Interaction · Computer Science 2025-05-26 Arnav Verma , Kushin Mukherjee , Christopher Potts , Elisa Kreiss , Judith E. Fan

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li