English
Related papers

Related papers: ComicScene154: A Scene Dataset for Comic Analysis

200 papers

We introduce a large dataset of narrative texts and questions about these texts, intended to be used in a machine comprehension task that requires reasoning using commonsense knowledge. Our dataset complements similar datasets in that we…

Computation and Language · Computer Science 2018-03-15 Simon Ostermann , Ashutosh Modi , Michael Roth , Stefan Thater , Manfred Pinkal

Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired readers, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Ragav Sachdeva , Andrew Zisserman

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Taewon Kang , Ming C. Lin

Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the images, layouts, and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Reshma Ramaprasad

When reading a literary piece, readers often make inferences about various characters' roles, personalities, relationships, intents, actions, etc. While humans can readily draw upon their past experiences to build such a character-centric…

Computation and Language · Computer Science 2021-09-14 Faeze Brahman , Meng Huang , Oyvind Tafjord , Chao Zhao , Mrinmaya Sachan , Snigdha Chaturvedi

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Cailin Zhuang , Ailin Huang , Yaoqi Hu , Jingwei Wu , Wei Cheng , Jiaqi Liao , Hongyuan Wang , Xinyao Liao , Weiwei Cai , Hengyuan Xu , Xuanyang Zhang , Xianfang Zeng , Zhewei Huang , Gang Yu , Chi Zhang

Text-to-image diffusion models have achieved high visual fidelity, yet precise control over scene semantics and fine-grained affective tone remains challenging. Human visual affect arises from the rapid integration of contextual meaning,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Li He , Longtai Zhang , Wenqiang Zhang , Yan Wang , Lizhe Qi

Understanding the semantics of visual scenes is a fundamental challenge in Computer Vision. A key aspect of this challenge is that objects sharing similar semantic meanings or functions can exhibit striking visual differences, making…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Rushikesh Zawar , Shaurya Dewan , Andrew F. Luo , Margaret M. Henderson , Michael J. Tarr , Leila Wehbe

Understanding another person's creative output requires a shared language of association. However, when training vision-language models such as CLIP, we rely on web-scraped datasets containing short, predominantly literal, alt-text. In this…

Computation and Language · Computer Science 2025-07-28 Ananya Sahu , Amith Ananthram , Kathleen McKeown

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natural images and high-resolution optical remote sensing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Lucrezia Tosato , Gianluca Lombardi , Ronny Hansch

Graphical novels such as comics and mangas are well known all over the world. The digital transition started to change the way people are reading comics, more and more on smartphones and tablets and less and less on paper. In the recent…

Multimedia · Computer Science 2018-04-17 Olivier Augereau , Motoi Iwata , Koichi Kise

Human perception of the world is shaped by a multitude of viewpoints and modalities. While many existing datasets focus on scene understanding from a certain perspective (e.g. egocentric or third-person views), our dataset offers a panoptic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Hao Chen , Yuqi Hou , Chenyuan Qu , Irene Testini , Xiaohan Hong , Jianbo Jiao

Today's research progress in the field of multi-document summarization is obstructed by the small number of available datasets. Since the acquisition of reference summaries is costly, existing datasets contain only hundreds of samples at…

Computation and Language · Computer Science 2020-02-18 Diego Antognini , Boi Faltings

Recent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Yi Zheng , Yifan Zhao , Mengyuan Ren , He Yan , Xiangju Lu , Junhui Liu , Jia Li

We position a narrative-centred computational model for high-level knowledge representation and reasoning in the context of a range of assistive technologies concerned with "visuo-spatial perception and cognition" tasks. Our proposed…

Artificial Intelligence · Computer Science 2013-06-25 Mehul Bhatt , Jakob Suchan , Carl Schultz

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-12 Shanshan Wang , Annamaria Mesaros , Toni Heittola , Tuomas Virtanen

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Zhan Shi , Xu Zhou , Xipeng Qiu , Xiaodan Zhu

Commonsense reasoning simulates the human ability to make presumptions about our physical world, and it is an indispensable cornerstone in building general AI systems. We propose a new commonsense reasoning dataset based on human's…

Artificial Intelligence · Computer Science 2020-10-21 Mo Yu , Xiaoxiao Guo , Yufei Feng , Xiaodan Zhu , Michael Greenspan , Murray Campbell

This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a necessary first stage for many downstream tasks like character…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Marc Serra Ortega , Emanuele Vivoli , Artemis Llabrés , Dimosthenis Karatzas