中文
相关论文

相关论文: From Panels to Prose: Generating Literary Narrativ…

200 篇论文

Creating engaging narratives from visual data is crucial for automated digital media consumption, assistive technologies, and interactive entertainment. This survey covers methodologies used in the generation of these narratives, focusing…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Daniel A. P. Oliveira , Eugénio Ribeiro , David Martins de Matos

Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering, and grounding, often in zero-shot settings. Comics…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Emanuele Vivoli , Mohamed Ali Souibgui , Andrey Barsky , Artemis LLabrés , Marco Bertini , Dimosthenis Karatzas

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (VLMs) have shown…

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence…

计算与语言 · 计算机科学 2025-08-21 Admitos Passadakis , Yingjin Song , Albert Gatt

Today, manga has gained worldwide popularity. However, the question of how various elements of manga, such as characters, text, and panel layouts, reflect the uniqueness of a particular work, or even define it, remains an unexplored area.…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Siyuan Feng , Teruya Yoshinaga , Katsuhiko Hayashi , Koki Washio , Hidetaka Kamigaito

We present a hierarchical knowledge graph framework for the structured semantic understanding of visual narratives, using comics as a representative domain for multimodal storytelling. The framework organizes narrative content across three…

多媒体 · 计算机科学 2025-11-18 Yi-Chun Chen

While manga is a popular entertainment form, creating manga is tedious, especially adding screentones to the created sketch, namely manga screening. Unfortunately, there is no existing method that tailors for automatic manga screening,…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jian Lin , Xueting Liu , Chengze Li , Minshan Xie , Tien-Tsin Wong

Manga, a widely celebrated Japanese comic art form, is renowned for its diverse narratives and distinct artistic styles. However, the inherently visual and intricate structure of Manga, which comprises images housing multiple panels, poses…

信息检索 · 计算机科学 2023-11-07 Conghao Tom Shen , Violet Yao , Yixin Liu

The process of adapting or repurposing manga pages is a time-consuming task that requires manga artists to manually work on every single screentone region and apply new patterns to create novel screentones across multiple panels. To address…

计算机视觉与模式识别 · 计算机科学 2023-06-08 Minshan Xie , Chengze Li , Tien-Tsin Wong

In the digital landscape, the ubiquity of data visualizations in media underscores the necessity for accessibility to ensure inclusivity for all users, including those with visual impairments. Current visual content often fails to cater to…

人机交互 · 计算机科学 2024-09-27 Qiang Xu , Thomas Hurtut

We present RaCig, a novel system for generating comic-style image sequences with consistent characters and expressive gestures. RaCig addresses two key challenges: (1) maintaining character identity and costume consistency across frames,…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yunhao Shui , Xuekuan Wang , Feng Qiu , Yuqiu Huang , Jinzhu Li , Haoyu Zheng , Jinru Han , Zhuo Zeng , Pengpeng Zhang , Jiarui Han , Keqiang Sun

We developed "Comicolorization", a semi-automatic colorization system for manga images. Given a monochrome manga and reference images as inputs, our system generates a plausible color version of the manga. This is the first work to address…

计算机视觉与模式识别 · 计算机科学 2017-09-29 Chie Furusawa , Kazuyuki Hiroshiba , Keisuke Ogaki , Yuri Odagiri

Drawing and annotating comic illustrations is a complex and difficult process. No existing machine learning algorithms have been developed to create comic illustrations based on descriptions of illustrations, or the dialogue in comics.…

计算机视觉与模式识别 · 计算机科学 2021-09-21 Ben Proven-Bessel , Zilong Zhao , Lydia Chen

A single image can convey a compelling story through logically connected visual clues, forming Chains-of-Reasoning (CoRs). We define these semantically rich images as Storytelling Images. By conveying multi-layered information that inspires…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Xiujie Song , Qi Jia , Shota Watanabe , Xiaoyi Pang , Ruijie Chen , Mengyue Wu , Kenny Q. Zhu

In this paper, we develop a MultiTask Learning (MTL) model to achieve dense predictions for comics panels to, in turn, facilitate the transfer of comics from one publication channel to another by assisting authors in the task of…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Deblina Bhattacharjee , Sabine Süsstrunk , Mathieu Salzmann

Chain-of-Thought reasoning has driven large language models to extend from thinking with text to thinking with images and videos. However, different modalities still have clear limitations: static images struggle to represent temporal…

人工智能 · 计算机科学 2026-02-04 Andong Chen , Wenxin Zhu , Qiuyu Ding , Yuchen Song , Muyun Yang , Tiejun Zhao

Can visual artworks created using generative visual algorithms inspire human creativity in storytelling? We asked writers to write creative stories from a starting prompt, and provided them with visuals created by generative AI models from…

人机交互 · 计算机科学 2021-10-29 Safinah Ali , Devi Parikh

Story visualization aims to generate a series of images that match the story described in texts, and it requires the generated images to satisfy high quality, alignment with the text description, and consistency in character identities.…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Wen Wang , Canyu Zhao , Hao Chen , Zhekai Chen , Kecheng Zheng , Chunhua Shen

We present a system to convert any set of images (e.g., a video clip or a photo album) into a storyboard. We aim to create multiple pleasing graphic representations of the content at interactive rates, so the user can explore and find the…

The comic domain is rapidly advancing with the development of single- and multi-page analysis and synthesis models. Recent benchmarks and datasets have been introduced to support and assess models' capabilities in tasks such as detection…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Emanuele Vivoli , Niccolò Biondi , Marco Bertini , Dimosthenis Karatzas