English
Related papers

Related papers: AnimeShooter: A Multi-Shot Animation Dataset for R…

200 papers

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given sounding event…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

We introduce TVStoryGen, a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved. Unlike other story generation datasets, TVStoryGen…

Computation and Language · Computer Science 2022-10-11 Mingda Chen , Kevin Gimpel

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Bingyuan Wang , Hengyu Meng , Zeyu Cai , Lanjiong Li , Yue Ma , Qifeng Chen , Zeyu Wang

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Yuhang Yang , Fan Zhang , Huaijin Pi , Shuai Guo , Guowei Xu , Wei Zhai , Yang Cao , Zheng-Jun Zha

The production of 2D animation follows an industry-standard workflow, encompassing four essential stages: character design, keyframe animation, in-betweening, and coloring. Our research focuses on reducing the labor costs in the above…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Yihao Meng , Hao Ouyang , Hanlin Wang , Qiuyu Wang , Wen Wang , Ka Leong Cheng , Zhiheng Liu , Yujun Shen , Huamin Qu

Recent advances in image generation, particularly diffusion models, have significantly lowered the barrier for creating sophisticated forgeries, making image manipulation detection and localization (IMDL) increasingly challenging. While…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Chenyang Zhu , Xing Zhang , Yuyang Sun , Ching-Chun Chang , Isao Echizen

Generative videos have the potential to revolutionize game development by autonomously creating new content. In this paper, we present GameFactory, a framework for action-controlled scene-generalizable game video generation. We first…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Jiwen Yu , Yiran Qin , Xintao Wang , Pengfei Wan , Di Zhang , Xihui Liu

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its…

Artificial Intelligence · Computer Science 2023-03-09 Yihan Cao , Siyu Li , Yixin Liu , Zhiling Yan , Yutong Dai , Philip S. Yu , Lichao Sun

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

It is a time-consuming and tedious work for manually colorizing anime line drawing images, which is an essential stage in cartoon animation creation pipeline. Reference-based line drawing colorization is a challenging task that relies on…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Yu Cao , Xiangqiao Meng , P. Y. Mok , Xueting Liu , Tong-Yee Lee , Ping Li

In the field of digital content creation, generating high-quality 3D characters from single images is challenging, especially given the complexities of various body poses and the issues of self-occlusion and pose ambiguity. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Hao-Yang Peng , Jia-Peng Zhang , Meng-Hao Guo , Yan-Pei Cao , Shi-Min Hu

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models. However, training a well-generalized multi-modal dialogue model remains challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Young-Jun Lee , Byungsoo Ko , Han-Gyu Kim , Jonghwan Hyeon , Ho-Jin Choi

Acquiring and annotating surgical data is often resource-intensive, ethical constraining, and requiring significant expert involvement. While generative AI models like text-to-image can alleviate data scarcity, incorporating spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Aditya Bhat , Rupak Bose , Chinedu Innocent Nwoye , Nicolas Padoy

Generative AI is reshaping art, gaming, and most notably animation. Recent breakthroughs in foundation and diffusion models have reduced the time and cost of producing animated content. Characters are central animation components, involving…

The gaming and entertainment industry is rapidly evolving, driven by immersive experiences and the integration of generative AI (GAI) technologies. Training such models effectively requires large-scale datasets that capture the diversity…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yuanzhi Li , Lebin Zhou , Nam Ling , Zhenghao Chen , Wei Wang , Wei Jiang

The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Zheng Zhan , Yushu Wu , Yifan Gong , Zichong Meng , Zhenglun Kong , Changdi Yang , Geng Yuan , Pu Zhao , Wei Niu , Yanzhi Wang

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on low-level visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Xiaoxue Wu , Xinyuan Chen , Yaohui Wang , Yu Qiao