中文
相关论文

相关论文: AnimeShooter: A Multi-Shot Animation Dataset for R…

200 篇论文

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given sounding event…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

We introduce TVStoryGen, a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved. Unlike other story generation datasets, TVStoryGen…

计算与语言 · 计算机科学 2022-10-11 Mingda Chen , Kevin Gimpel

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang

Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Bingyuan Wang , Hengyu Meng , Zeyu Cai , Lanjiong Li , Yue Ma , Qifeng Chen , Zeyu Wang

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Yuhang Yang , Fan Zhang , Huaijin Pi , Shuai Guo , Guowei Xu , Wei Zhai , Yang Cao , Zheng-Jun Zha

The production of 2D animation follows an industry-standard workflow, encompassing four essential stages: character design, keyframe animation, in-betweening, and coloring. Our research focuses on reducing the labor costs in the above…

计算机视觉与模式识别 · 计算机科学 2025-01-31 Yihao Meng , Hao Ouyang , Hanlin Wang , Qiuyu Wang , Wen Wang , Ka Leong Cheng , Zhiheng Liu , Yujun Shen , Huamin Qu

Recent advances in image generation, particularly diffusion models, have significantly lowered the barrier for creating sophisticated forgeries, making image manipulation detection and localization (IMDL) increasingly challenging. While…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Chenyang Zhu , Xing Zhang , Yuyang Sun , Ching-Chun Chang , Isao Echizen

Generative videos have the potential to revolutionize game development by autonomously creating new content. In this paper, we present GameFactory, a framework for action-controlled scene-generalizable game video generation. We first…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Jiwen Yu , Yiran Qin , Xintao Wang , Pengfei Wan , Di Zhang , Xihui Liu

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its…

人工智能 · 计算机科学 2023-03-09 Yihan Cao , Siyu Li , Yixin Liu , Zhiling Yan , Yutong Dai , Philip S. Yu , Lichao Sun

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

声音 · 计算机科学 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

It is a time-consuming and tedious work for manually colorizing anime line drawing images, which is an essential stage in cartoon animation creation pipeline. Reference-based line drawing colorization is a challenging task that relies on…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Yu Cao , Xiangqiao Meng , P. Y. Mok , Xueting Liu , Tong-Yee Lee , Ping Li

In the field of digital content creation, generating high-quality 3D characters from single images is challenging, especially given the complexities of various body poses and the issues of self-occlusion and pose ambiguity. In this paper,…

计算机视觉与模式识别 · 计算机科学 2024-07-11 Hao-Yang Peng , Jia-Peng Zhang , Meng-Hao Guo , Yan-Pei Cao , Shi-Min Hu

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models. However, training a well-generalized multi-modal dialogue model remains challenging due to…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Young-Jun Lee , Byungsoo Ko , Han-Gyu Kim , Jonghwan Hyeon , Ho-Jin Choi

Acquiring and annotating surgical data is often resource-intensive, ethical constraining, and requiring significant expert involvement. While generative AI models like text-to-image can alleviate data scarcity, incorporating spatial…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Aditya Bhat , Rupak Bose , Chinedu Innocent Nwoye , Nicolas Padoy

Generative AI is reshaping art, gaming, and most notably animation. Recent breakthroughs in foundation and diffusion models have reduced the time and cost of producing animated content. Characters are central animation components, involving…

The gaming and entertainment industry is rapidly evolving, driven by immersive experiences and the integration of generative AI (GAI) technologies. Training such models effectively requires large-scale datasets that capture the diversity…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yuanzhi Li , Lebin Zhou , Nam Ling , Zhenghao Chen , Wei Wang , Wei Jiang

The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Zheng Zhan , Yushu Wu , Yifan Gong , Zichong Meng , Zhenglun Kong , Changdi Yang , Geng Yuan , Pu Zhao , Wei Niu , Yanzhi Wang

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on low-level visual…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Xiaoxue Wu , Xinyuan Chen , Yaohui Wang , Yu Qiao