English
Related papers

Related papers: Training-free Camera Control for Video Generation

200 papers

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Mingzhi Sheng , Zekai Gu , Peng Li , Cheng Lin , Hao-Xiang Guo , Ying-Cong Chen , Yuan Liu

We propose a novel inference technique based on a pretrained diffusion model for text-conditional video generation. Our approach, called FIFO-Diffusion, is conceptually capable of generating infinitely long videos without additional…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Jihwan Kim , Junoh Kang , Jinyoung Choi , Bohyung Han

Long-trajectory video generation is a crucial yet challenging task for world modeling primarily due to the limited scalability of existing video diffusion models (VDMs). Autoregressive models, while offering infinite rollout, suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Junyi Ouyang , Wenbin Teng , Gonglin Chen , Yajie Zhao , Haiwei Chen

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving physical consistency…

This paper introduces Camera-free Diffusion (CamFreeDiff) model for 360-degree image outpainting from a single camera-free image and text description. This method distinguishes itself from existing strategies, such as MVDiffusion, by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Xiaoding Yuan , Shitao Tang , Kejie Li , Alan Yuille , Peng Wang

We propose FreeSim, a camera simulation method for autonomous driving. FreeSim emphasizes high-quality rendering from viewpoints beyond the recorded ego trajectories. In such viewpoints, previous methods have unacceptable degradation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Lue Fan , Hao Zhang , Qitai Wang , Hongsheng Li , Zhaoxiang Zhang

Current large-scale generative models have impressive efficiency in generating high-quality images based on text prompts. However, they lack the ability to precisely control the size and position of objects in the generated image. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Jiafeng Mao , Xueting Wang

We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes each point with a segmentation ID, a temporally consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Zhiyuan Zhang , Can Wang , Dongdong Chen , Jing Liao

In this paper, we present VideoGen, a text-to-video generation approach, which can generate a high-definition video with high frame fidelity and strong temporal consistency using reference-guided latent diffusion. We leverage an…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Xin Li , Wenqing Chu , Ye Wu , Weihang Yuan , Fanglong Liu , Qi Zhang , Fu Li , Haocheng Feng , Errui Ding , Jingdong Wang

Large-scale generative models have achieved remarkable advancements in various visual tasks, yet their application to shadow removal in images remains challenging. These models often generate diverse, realistic details without adequate…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Xinjie Li , Yang Zhao , Dong Wang , Yuan Chen , Li Cao , Xiaoping Liu

Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video generation treat the video as a whole, neglecting intricate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Yifan Shen , Peiyuan Zhu , Zijian Li , Shaoan Xie , Namrata Deka , Zongfang Liu , Zeyu Tang , Guangyi Chen , Kun Zhang

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts, subject references,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Yusuf Dalva , Guocheng Gordon Qian , Maya Goldenberg , Tsai-Shien Chen , Kfir Aberman , Sergey Tulyakov , Pinar Yanardag , Kuan-Chieh Jackson Wang

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

We present DiffIR2VR-Zero, a zero-shot framework that enables any pre-trained image restoration diffusion model to perform high-quality video restoration without additional training. While image diffusion models have shown remarkable…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Chang-Han Yeh , Hau-Shiang Shiu , Chin-Yang Lin , Zhixiang Wang , Chi-Wei Hsiao , Ting-Hsuan Chen , Yu-Lun Liu

Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Avinash Paliwal , Adithya Iyer , Shivin Yadav , Muhammad Ali Afridi , Midhun Harikumar

Diffusion models have revolutionized image generation in recent years, yet they are still limited to a few sizes and aspect ratios. We propose ElasticDiffusion, a novel training-free decoding method that enables pretrained text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Moayed Haji-Ali , Guha Balakrishnan , Vicente Ordonez

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Eyal Molad , Eliahu Horwitz , Dani Valevski , Alex Rav Acha , Yossi Matias , Yael Pritch , Yaniv Leviathan , Yedid Hoshen

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Zuhao Yang , Jiahui Zhang , Yingchen Yu , Shijian Lu , Song Bai

We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Ruihang Chu , Yefei He , Zhekai Chen , Shiwei Zhang , Xiaogang Xu , Bin Xia , Dingdong Wang , Hongwei Yi , Xihui Liu , Hengshuang Zhao , Yu Liu , Yingya Zhang , Yujiu Yang
‹ Prev 1 8 9 10 Next ›