中文
相关论文

相关论文: Panel-by-Panel Souls: A Performative Workflow for …

200 篇论文

Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing and lack the…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Chenxi Li

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Hikaru Ikuta , Leslie Wöhler , Kiyoharu Aizawa

Instruction-guided image editing offers an intuitive way for users to edit images with natural language. However, diffusion-based editing models often struggle to accurately interpret complex user instructions, especially those involving…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Ziyun Zeng , Hang Hua , Jiebo Luo

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding and underdeveloped…

计算与语言 · 计算机科学 2024-12-06 Philip Lippmann , Konrad Skublicki , Joshua Tanner , Shonosuke Ishiwatari , Jie Yang

Generative AI has made image creation more accessible, yet aligning outputs with nuanced creative intent remains challenging, particularly for non-experts. Existing tools often require users to externalize ideas through prompts or…

人机交互 · 计算机科学 2025-08-11 Daniel Lee , Nikhil Sharma , Donghoon Shin , DaEun Choi , Harsh Sharma , Jeonghwan Kim , Heng Ji

This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these challenges, we propose…

声音 · 计算机科学 2024-10-15 Megha Sharma , Muhammad Taimoor Haseeb , Gus Xia , Yoshimasa Tsuruoka

Generative AI is reshaping knowledge work, yet existing research focuses predominantly on software engineering and the natural sciences, with limited methodological exploration for the humanities and social sciences. Positioned as a…

人工智能 · 计算机科学 2026-02-20 Yi-Chih Huang

Inspired by how the human brain employs a higher number of neural pathways when describing a highly focused subject, we show that deep attentive models used for the main vision-language task of image captioning, could be extended to achieve…

计算机视觉与模式识别 · 计算机科学 2021-09-01 Zanyar Zohourianshahzadi , Jugal K. Kalita

The development of AI-driven generative audio mirrors broader AI trends, often prioritizing immediate accessibility at the expense of explainability. Consequently, integrating such tools into sustained artistic practice remains a…

声音 · 计算机科学 2024-07-23 Austin Tecks , Thomas Peschlow , Gabriel Vigliensoni

The topic of facial landmark detection has been widely covered for pictures of human faces, but it is still a challenge for drawings. Indeed, the proportions and symmetry of standard human faces are not always used for comics or mangas. The…

计算机视觉与模式识别 · 计算机科学 2018-11-09 Marco Stricker , Olivier Augereau , Koichi Kise , Motoi Iwata

We introduce the concept of "Design Agents" for engineering applications, particularly focusing on the automotive design process, while emphasizing that our approach can be readily extended to other engineering and design domains. Our…

人工智能 · 计算机科学 2025-12-04 Mohamed Elrefaie , Janet Qian , Raina Wu , Qian Chen , Angela Dai , Faez Ahmed

Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features, which has attracted considerable attention from both natural language processing and computer vision communities. Currently, most comics…

计算与语言 · 计算机科学 2023-10-27 Hongcheng Guo , Boyang Wang , Jiaqi Bai , Jiaheng Liu , Jian Yang , Zhoujun Li

Visual blends combine elements from two distinct visual concepts into a single, integrated image, with the goal of conveying ideas through imaginative and often thought-provoking visuals. Communicating abstract concepts through visual…

人机交互 · 计算机科学 2025-02-25 Zhida Sun , Zhenyao Zhang , Yue Zhang , Min Lu , Dani Lischinski , Daniel Cohen-Or , Hui Huang

This paper proposes a framework for computational modeling of artistic painting algorithms, inspired by human creative practices. Based on examples from expert artists and from the author's own experience, the paper argues that creative…

人工智能 · 计算机科学 2022-05-24 Aaron Hertzmann

Drawing supports learning by externalizing mental models, but providing timely feedback at scale remains challenging. We present Draw2Learn, a system that explores how AI can act as a supportive teammate during drawing-based learning. The…

人机交互 · 计算机科学 2026-02-03 Yuqi Hang

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic…

多媒体 · 计算机科学 2022-04-07 Liyang Chen , Zhiyong Wu , Jun Ling , Runnan Li , Xu Tan , Sheng Zhao

Professional designers work from client briefs that specify goals and constraints but often lack concrete design details. Translating these abstract requirements into visual designs poses a central challenge, yet existing tools address…

人机交互 · 计算机科学 2026-04-14 Kotaro Kikuchi , Nami Ogawa

This thesis presents an innovative approach to automate video thumbnail selection for traditional broadcast content. Our methodology establishes stringent criteria for diverse, representative, and aesthetically pleasing thumbnails,…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Elia Fantini

While the increased integration of AI technologies into interactive systems enables them to solve an equally increasing number of tasks, the black box problem of AI models continues to spread throughout the interactive system as a whole.…

人机交互 · 计算机科学 2025-06-04 Sebe Vanbrabant , Gustavo Rovelo Ruiz , Davy Vanacken

People often imagine relevant scenes to aid in the writing process. In this work, we aim to utilize visual information for composition in the same manner as humans. We propose a method, LIVE, that makes pre-trained language models (PLMs)…

计算与语言 · 计算机科学 2023-06-16 Tianyi Tang , Yushuo Chen , Yifan Du , Junyi Li , Wayne Xin Zhao , Ji-Rong Wen