English
Related papers

Related papers: RealCam-I2V: Real-World Image-to-Video Generation …

200 papers

Recent advancements have integrated camera pose as a user-friendly and physics-informed condition in video diffusion models, enabling precise camera control. In this paper, we identify one of the key challenges as effectively modeling noisy…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Guangcong Zheng , Teng Li , Rui Jiang , Yehao Lu , Tao Wu , Xi Li

Image-to-video (I2V) generation tasks always suffer from keeping high fidelity in the open domains. Traditional image animation techniques primarily focus on specific domains such as faces or human poses, making them difficult to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Weijie Li , Litong Gong , Yiran Zhu , Fanda Fan , Biao Wang , Tiezheng Ge , Bo Zheng

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Min-Jung Kim , Jeongho Kim , Hoiyeong Jin , Junha Hyung , Jaegul Choo

The field of generative models has recently witnessed significant progress, with diffusion models showing remarkable performance in image generation. In light of this success, there is a growing interest in exploring the application of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ariel Lapid , Idan Achituve , Lior Bracha , Ethan Fetaya

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ashkan Taghipour , Morteza Ghahremani , Mohammed Bennamoun , Farid Boussaid , Aref Miri Rekavandi , Zinuo Li , Qiuhong Ke , Hamid Laga

Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. In content creation workflows, precise and simultaneous control over camera motion, object motion, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Sixiao Zheng , Zimian Peng , Yanpeng Zhou , Yi Zhu , Hang Xu , Xiangru Huang , Yanwei Fu

Video generation has increasingly gained interest in both academia and industry. Although commercial tools can generate plausible videos, there is a limited number of open-source models available for researchers and engineers. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Haoxin Chen , Menghan Xia , Yingqing He , Yong Zhang , Xiaodong Cun , Shaoshu Yang , Jinbo Xing , Yaofang Liu , Qifeng Chen , Xintao Wang , Chao Weng , Ying Shan

Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Yaosi Hu , Chong Luo , Zhenzhong Chen

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring the urgent need for effective digital watermarking. While existing watermarking methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Guanjie Wang , Zehua Ma , Han Fang , Weiming Zhang

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xingrui Wang , Xin Li , Yaosi Hu , Hanxin Zhu , Chen Hou , Cuiling Lan , Zhibo Chen

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Mingyang Wu , Ashirbad Mishra , Soumik Dey , Shuo Xing , Naveen Ravipati , Hansi Wu , Binbin Li , Zhengzhong Tu

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Hao He , Ceyuan Yang , Shanchuan Lin , Yinghao Xu , Meng Wei , Liangke Gui , Qi Zhao , Gordon Wetzstein , Lu Jiang , Hongsheng Li

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Songcen Xu , Hang Xu , Xiaodan Liang

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

Video-to-video synthesis (vid2vid) aims for converting high-level semantic inputs to photorealistic videos. While existing vid2vid methods can achieve short-term temporal consistency, they fail to ensure the long-term one. This is because…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Arun Mallya , Ting-Chun Wang , Karan Sapra , Ming-Yu Liu

AI video generation has lowered barriers to video creation, but current tools still struggle with inconsistency. Filmmakers often find that clips fail to match characters and backgrounds, making it difficult to build coherent sequences. A…

Human-Computer Interaction · Computer Science 2025-12-22 Hye-Young Jo , Mose Sakashita , Aditi Mishra , Ryo Suzuki , Koichiro Niinuma , Aakar Gupta

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Qiucheng Wu , Handong Zhao , Zhixin Shu , Jing Shi , Yang Zhang , Shiyu Chang