中文
相关论文

相关论文: The Manga Whisperer: Automatically Generating Tran…

200 篇论文

A big part of achieving Artificial General Intelligence(AGI) is to build a machine that can see and listen like humans. Much work has focused on designing models for image classification, video classification, object detection, pose…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Ruotian Luo

The goal of this paper is automatic character-aware subtitle generation. Given a video and a minimal amount of metadata, we propose an audio-visual method that generates a full transcript of the dialogue, with precise speech timestamps, and…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Bruno Korbar , Jaesung Huh , Andrew Zisserman

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding and underdeveloped…

计算与语言 · 计算机科学 2024-12-06 Philip Lippmann , Konrad Skublicki , Joshua Tanner , Shonosuke Ishiwatari , Jie Yang

Speaker diarization, usually denoted as the ''who spoke when'' task, turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, sound effects...).…

多媒体 · 计算机科学 2019-04-22 Xavier Bost , Georges Linares

Japanese comics (called manga) are traditionally created in monochrome format. In recent years, in addition to monochrome comics, full color comics, a more attractive medium, have appeared. Unfortunately, color comics require manual…

计算机视觉与模式识别 · 计算机科学 2021-07-19 Yugo Shimizu , Ryosuke Furuta , Delong Ouyang , Yukinobu Taniguchi , Ryota Hinami , Shonosuke Ishiwatari

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on…

计算与语言 · 计算机科学 2026-01-27 Jeonghun Baek , Kazuki Egashira , Shota Onohara , Atsuyuki Miyai , Yuki Imajuku , Hikaru Ikuta , Kiyoharu Aizawa

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Hikaru Ikuta , Leslie Wöhler , Kiyoharu Aizawa

Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker. Particularly in the presence of overlapping or noisy speech, these…

音频与语音处理 · 电气工程与系统科学 2024-06-06 Christoph Boeddeker , Tobias Cord-Landwehr , Reinhold Haeb-Umbach

The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small-scale or single-style test sets. We introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Emanuele Vivoli , Marco Bertini , Dimosthenis Karatzas

Onomatopoeia is an important element for textual messaging in manga. Unlike character dialogue in manga, onomatopoeic expressions are visually stylized, with variations in shape, size, and placement that reflect the scene's intensity and…

多媒体 · 计算机科学 2025-09-30 Takara Taniguchi , Wataru Shimoda , Kota Yamaguchi , Hideki Nakayama

Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the "gutters" between panels. To follow the…

计算机视觉与模式识别 · 计算机科学 2017-05-09 Mohit Iyyer , Varun Manjunatha , Anupam Guha , Yogarshi Vyas , Jordan Boyd-Graber , Hal Daumé , Larry Davis

While manga is a popular entertainment form, creating manga is tedious, especially adding screentones to the created sketch, namely manga screening. Unfortunately, there is no existing method that tailors for automatic manga screening,…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jian Lin , Xueting Liu , Chengze Li , Minshan Xie , Tien-Tsin Wong

This work studies the task of glossification, of which the aim is to em transcribe natural spoken language sentences for the Deaf (hard-of-hearing) community to ordered sign language glosses. Previous sequence-to-sequence language models…

计算与语言 · 计算机科学 2021-12-20 Dongxu Li , Chenchen Xu , Liu Liu , Yiran Zhong , Rong Wang , Lars Petersson , Hongdong Li

Recent work in computer vision has yielded impressive results in automatically describing images with natural language. Most of these systems generate captions in a sin- gle language, requiring multiple language-specific models to build a…

计算机视觉与模式识别 · 计算机科学 2017-06-21 Satoshi Tsutsui , David Crandall

Comic strips are a popular and expressive form of visual storytelling that can convey humor, emotion, and information. However, they are inaccessible to the BLV (Blind or Low Vision) community, who cannot perceive the images, layouts, and…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Reshma Ramaprasad

The Japanese comic format known as Manga is popular all over the world. It is traditionally produced in black and white, and colorization is time consuming and costly. Automatic colorization methods generally rely on greyscale values, which…

图形学 · 计算机科学 2017-06-22 Paulina Hensman , Kiyoharu Aizawa

Manga, or comics, which are a type of multimodal artwork, have been left behind in the recent trend of deep learning applications because of the lack of a proper dataset. Hence, we built Manga109, a dataset consisting of a variety of 109…

多媒体 · 计算机科学 2020-05-13 Kiyoharu Aizawa , Azuma Fujimoto , Atsushi Otsubo , Toru Ogawa , Yusuke Matsui , Koki Tsubota , Hikaru Ikuta

As generative AI becomes more prevalent, it is important to study how human users interact with such models. In this work, we investigate how people use text-to-image models to generate desired target images. To study this interaction, we…

人工智能 · 计算机科学 2024-06-18 Kailas Vodrahalli , James Zou

Current text-to-image models struggle to render the nuanced facial expressions required for compelling manga narratives, largely due to the ambiguity of language itself. To bridge this gap, we introduce an interactive system built on a…

人机交互 · 计算机科学 2025-11-21 Qing Zhang , Jing Huang , Yifei Huang , Jun Rekimoto

End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Muyao Wang , Zeke Xie , Yanhao Chen , Lixin Xiu , Hideki Nakayama