中文
相关论文

相关论文: MangaFlow: An End-to-End Agentic Framework for Con…

200 篇论文

Generating comics through text is widely studied. However, there are few studies on generating multi-panel Manga (Japanese comics) solely based on plain text. Japanese manga contains multiple panels on a single page, with characteristics…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Siyu Chen , Dengjie Li , Zenghao Bao , Yao Zhou , Lingfeng Tan , Yujie Zhong , Zheng Zhao

Despite the promise of autonomous agentic reasoning, existing workflow generation methods frequently produce fragile, unexecutable plans due to unconstrained LLM-driven construction. We introduce MermaidFlow, a framework that redefines the…

Generative art unlocks boundless creative possibilities, yet its full potential remains untapped due to the technical expertise required for advanced architectural concepts and computational workflows. To bridge this gap, we present…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Duc-Hung Nguyen , Huu-Phuc Huynh , Minh-Triet Tran , Trung-Nghia Le

Current text-to-image models struggle to render the nuanced facial expressions required for compelling manga narratives, largely due to the ambiguity of language itself. To bridge this gap, we introduce an interactive system built on a…

人机交互 · 计算机科学 2025-11-21 Qing Zhang , Jing Huang , Yifei Huang , Jun Rekimoto

Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Kiymet Akdemir , Tahira Kazimi , Pinar Yanardag

We present AniME, a director-oriented multi-agent system for automated long-form anime production, covering the full workflow from a story to the final video. The director agent keeps a global memory for the whole workflow, and coordinates…

Agent systems based on large language models (LLMs) have shown great potential in complex reasoning tasks, but building efficient and generalizable workflows remains a major challenge. Most existing approaches rely on manually designed…

Analog/mixed-signal circuits are key for interfacing electronics with the physical world. Their design, however, remains a largely handcrafted process, resulting in long and error-prone design cycles. While the recent rise of AI-based…

机器学习 · 计算机科学 2026-01-15 Mohsen Ahmadzadeh , Kaichang Chen , Georges Gielen

Understanding how humans communicate and perceive narratives is important for media technology research and development. This is particularly important in current times when there are tools and algorithms that are easily available for…

人工智能 · 计算机科学 2023-12-15 Yi-Chun Chen , Arnav Jhala

The core challenge for streaming video generation is maintaining the content consistency in long context, which poses high requirement for the memory design. Most existing solutions maintain the memory by compressing historical frames with…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Sihui Ji , Xi Chen , Shuai Yang , Xin Tao , Pengfei Wan , Hengshuang Zhao

While current AI illustration tools can generate high-quality images from text prompts, they rarely reveal the step-by-step procedure that human artists follow. We present SakugaFlow, a four-stage pipeline that pairs diffusion-based image…

人机交互 · 计算机科学 2025-06-11 Kazuki Kawamura , Jun Rekimoto

The process of adapting or repurposing manga pages is a time-consuming task that requires manga artists to manually work on every single screentone region and apply new patterns to create novel screentones across multiple panels. To address…

计算机视觉与模式识别 · 计算机科学 2023-06-08 Minshan Xie , Chengze Li , Tien-Tsin Wong

Recent generative image editing methods adopt layered representations to mitigate the entangled nature of raster images and improve controllability, typically relying on object-based segmentation. However, such strategies may fail to…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Tianyu Zhang , Dongchi Li , Keiichi Sawada , Haoran Xie

This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these challenges, we propose…

声音 · 计算机科学 2024-10-15 Megha Sharma , Muhammad Taimoor Haseeb , Gus Xia , Yoshimasa Tsuruoka

Creating an animated data video enriched with audio narration takes a significant amount of time and effort and requires expertise. Users not only need to design complex animations, but also turn written text scripts into audio narrations…

人机交互 · 计算机科学 2024-06-10 Yun Wang , Leixian Shen , Zhengxin You , Xinhuan Shu , Bongshin Lee , John Thompson , Haidong Zhang , Dongmei Zhang

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Bingyuan Wang , Hengyu Meng , Zeyu Cai , Lanjiong Li , Yue Ma , Qifeng Chen , Zeyu Wang

Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocation capabilities. Yet, existing agents remain at the mercy of underlying black-box image…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Junyan Ye , Jun He , Zilong Huang , Dongzhi Jiang , Xuan Yang , Rui Chen , Weijia Li

Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired readers, limiting…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Ragav Sachdeva , Andrew Zisserman

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Hikaru Ikuta , Leslie Wöhler , Kiyoharu Aizawa

Formulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with…

计算机视觉与模式识别 · 计算机科学 2022-05-16 Xintian Wu , Qihang Zhang , Yiming Wu , Huanyu Wang , Songyuan Li , Lingyun Sun , Xi Li
‹ 上一页 1 2 3 10 下一页 ›