中文
相关论文

相关论文: The Manga Whisperer: Automatically Generating Tran…

200 篇论文

In this paper, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes by analyzing the…

计算机视觉与模式识别 · 计算机科学 2021-01-28 Xin Yang , Zongliang Ma , Letian Yu , Ying Cao , Baocai Yin , Xiaopeng Wei , Qiang Zhang , Rynson W. H. Lau

This study extracted and analyzed the linguistic speech patterns that characterize Japanese anime or game characters. Conventional morphological analyzers, such as MeCab, segment words with high performance, but they are unable to segment…

计算与语言 · 计算机科学 2022-03-08 Mika Kishino , Kanako Komiya

With the blooming of various Pre-trained Language Models (PLMs), Machine Reading Comprehension (MRC) has embraced significant improvements on various benchmarks and even surpass human performances. However, the existing works only target on…

计算与语言 · 计算机科学 2020-11-16 Yiming Cui , Ting Liu , Shijin Wang , Guoping Hu

Automatic high-quality rendering of anime scenes from complex real-world images is of significant practical value. The challenges of this task lie in the complexity of the scenes, the unique features of anime style, and the lack of…

计算机视觉与模式识别 · 计算机科学 2023-08-25 Yuxin Jiang , Liming Jiang , Shuai Yang , Chen Change Loy

Greyscale image colorization for applications in image restoration has seen significant improvements in recent years. Many of these techniques that use learning-based methods struggle to effectively colorize sparse inputs. With the…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Harrish Thasarathan , Kamyar Nazeri , Mehran Ebrahimi

Journaling can potentially serve as an effective method for autistic adolescents to improve narrative skills. However, its text-centric nature and high executive functioning demands present barriers to practice. We present Autiverse, an…

人机交互 · 计算机科学 2026-01-27 Migyeong Yang , Kyungah Lee , Jinyoung Han , SoHyun Park , Young-Ho Kim

Speaker diarization (SD) is the task of answering "who spoke when" in a multi-speaker audio stream. Classically, an SD system clusters segments of speech belonging to an individual speaker's identity. Recent years have seen substantial…

音频与语音处理 · 电气工程与系统科学 2026-04-24 Nikhil Raghav

The topic of facial landmark detection has been widely covered for pictures of human faces, but it is still a challenge for drawings. Indeed, the proportions and symmetry of standard human faces are not always used for comics or mangas. The…

计算机视觉与模式识别 · 计算机科学 2018-11-09 Marco Stricker , Olivier Augereau , Koichi Kise , Motoi Iwata

Dialogue summarization aims to condense the original dialogue into a shorter version covering salient information, which is a crucial way to reduce dialogue data overload. Recently, the promising achievements in both dialogue systems and…

计算与语言 · 计算机科学 2022-04-29 Xiachong Feng , Xiaocheng Feng , Bing Qin

Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations. Currently, most semantic parsing methods are not able to utilize contextual information (e.g. dialogue and comments…

计算与语言 · 计算机科学 2020-11-03 Zhuang Li , Lizhen Qu , Gholamreza Haffari

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech…

计算与语言 · 计算机科学 2023-12-06 Yihan Wu , Junliang Guo , Xu Tan , Chen Zhang , Bohan Li , Ruihua Song , Lei He , Sheng Zhao , Arul Menezes , Jiang Bian

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

计算机视觉与模式识别 · 计算机科学 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

Human conversation is a complex mechanism with subtle nuances. It is hence an ambitious goal to develop artificial intelligence agents that can participate fluently in a conversation. While we are still far from achieving this goal, recent…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Unnat Jain , Svetlana Lazebnik , Alexander Schwing

We present an inference-time adaptation method that tailors a pretrained image editing model to each input manga image using only the input image itself. Despite recent progress in pretrained image editing, such models often underperform on…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Ryosuke Furuta

We explore the use of long-context capabilities in large language models to create synthetic reading comprehension data from entire books. Previous efforts to construct such datasets relied on crowd-sourcing, but the emergence of…

Frequently Asked Questions (FAQs) refer to the most common inquiries about specific content. They serve as content comprehension aids by simplifying topics and enhancing understanding through succinct presentation of information. In this…

计算与语言 · 计算机科学 2024-11-20 Sahil Kale , Gautam Khaire , Jay Patankar

Vocabulary learning support tools have widely exploited existing materials, e.g., stories or video clips, as contexts to help users memorize each target word. However, these tools could not provide a coherent context for any target words of…

人机交互 · 计算机科学 2023-11-01 Zhenhui Peng , Xingbo Wang , Qiushi Han , Junkai Zhu , Xiaojuan Ma , Huamin Qu

The process of generating fully colorized drawings from sketches is a large, usually costly bottleneck in the manga and anime industry. In this study, we examine multiple models for image-to-image translation between anime characters and…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Tai Vu , Robert Yang

Today, manga has gained worldwide popularity. However, the question of how various elements of manga, such as characters, text, and panel layouts, reflect the uniqueness of a particular work, or even define it, remains an unexplored area.…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Siyuan Feng , Teruya Yoshinaga , Katsuhiko Hayashi , Koki Washio , Hidetaka Kamigaito

In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show an interesting finding that while Whisper is very robust…

声音 · 计算机科学 2023-10-10 Yuan Gong , Sameer Khurana , Leonid Karlinsky , James Glass
‹ 上一页 1 8 9 10 下一页 ›