中文
相关论文

相关论文: AniMatrix: An Anime Video Generation Model that Th…

200 篇论文

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively,…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Saman Motamed , Laura Culp , Kevin Swersky , Priyank Jaini , Robert Geirhos

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

Video generation models are increasingly used as world simulators for storytelling, simulation, and embodied AI. As these models advance, a key question arises: do generated videos obey the physical laws of the real world? Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Qin Zhang , Peiyu Jing , Hong-Xing Yu , Fangqiang Ding , Fan Nie , Weimin Wang , Yilun Du , James Zou , Jiajun Wu , Bing Shuai

We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editing conditions. Instead of relying on large-scale vision-language models or per-subject…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Yixuan Han

Despite significant advances in video synthesis, research into multi-shot video generation remains in its infancy. Even with scaled-up models and massive datasets, the shot transition capabilities remain rudimentary and unstable, largely…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Xiaoxue Wu , Bingjie Gao , Yu Qiao , Yaohui Wang , Xinyuan Chen

Incorporating camera intrinsics into video generation models offers a principled way to control not only scene dynamics but also the imaging process that governs visual appearance. Prior work has primarily focused on extrinsic control, such…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Debabrata Mandal , Zhihan Peng , Yujie Wang , Praneeth Chakravarthula

Interactive motion synthesis is essential in creating immersive experiences in entertainment applications, such as video games and virtual reality. However, generating animations that are both high-quality and contextually responsive…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Tianyu Li , Calvin Qiao , Guanqiao Ren , KangKang Yin , Sehoon Ha

Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an image to generate a specific motion trajectory desired by…

计算机视觉与模式识别 · 计算机科学 2022-03-08 Ligong Han , Jian Ren , Hsin-Ying Lee , Francesco Barbieri , Kyle Olszewski , Shervin Minaee , Dimitris Metaxas , Sergey Tulyakov

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhexiao Xiong , Yizhi Song , Liu He , Wei Xiong , Yu Yuan , Feng Qiao , Nathan Jacobs

Real-time in-between motion generation is universally required in games and highly desirable in existing animation pipelines. Its core challenge lies in the need to satisfy three critical conditions simultaneously: quality, controllability…

图形学 · 计算机科学 2022-05-06 Xiangjun Tang , He Wang , Bo Hu , Xu Gong , Ruifan Yi , Qilong Kou , Xiaogang Jin

We introduce Matrix-Game, an interactive world foundation model for controllable game world generation. Matrix-Game is trained using a two-stage pipeline that first performs large-scale unlabeled pretraining for environment understanding,…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yifan Zhang , Chunli Peng , Boyang Wang , Puyi Wang , Qingcheng Zhu , Fei Kang , Biao Jiang , Zedong Gao , Eric Li , Yang Liu , Yahui Zhou

Recent diffusion-based video generation models can synthesize visually plausible videos, yet they often struggle to satisfy physical constraints. A key reason is that most existing approaches remain single-stage: they entangle high-level…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Yibo Zhao , Hengjia Li , Xiaofei He , Boxi Wu

Generating controllable character animation from a reference image and motion guidance remains a challenging task due to the inherent difficulty of injecting appearance and motion cues into video diffusion models. Prior works often rely on…

图形学 · 计算机科学 2025-07-03 Guian Fang , Yuchao Gu , Mike Zheng Shou

The production of 2D animation follows an industry-standard workflow, encompassing four essential stages: character design, keyframe animation, in-betweening, and coloring. Our research focuses on reducing the labor costs in the above…

计算机视觉与模式识别 · 计算机科学 2025-01-31 Yihao Meng , Hao Ouyang , Hanlin Wang , Qiuyu Wang , Wen Wang , Ka Leong Cheng , Zhiheng Liu , Yujun Shen , Huamin Qu

Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when "anti-aesthetic" outputs are requested for artistic or critical purposes. This adherence prioritizes…

计算机与社会 · 计算机科学 2026-05-13 Wenqi Marshall Guo , Qingyun Qian , Khalad Hasan , Shan Du

Learning an animatable and clothed human avatar model with vivid dynamics and photorealistic appearance from multi-view videos is an important foundational research problem in computer graphics and vision. Fueled by recent advances in…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Heming Zhu , Guoxing Sun , Christian Theobalt , Marc Habermann

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zongyang Qiu , Bingyuan Wang , Xingbei Chen , Yingqing He , Zeyu Wang

End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Muyao Wang , Zeke Xie , Yanhao Chen , Lixin Xiu , Hideki Nakayama

Recently, video generation has achieved significant rapid development based on superior text-to-image generation techniques. In this work, we propose a high fidelity framework for image-to-video generation, named AtomoVideo. Based on…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Litong Gong , Yiran Zhu , Weijie Li , Xiaoyang Kang , Biao Wang , Tiezheng Ge , Bo Zheng