中文
相关论文

相关论文: Zero-Shot Personalized Camera Motion Control for I…

200 篇论文

We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combines it with a Stable…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Zitian Zhang , Frédéric Fortier-Chouinard , Mathieu Garon , Anand Bhattad , Jean-François Lalonde

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Min-Jung Kim , Jeongho Kim , Hoiyeong Jin , Junha Hyung , Jaegul Choo

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However,…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Wanquan Feng , Jiawei Liu , Pengqi Tu , Tianhao Qi , Mingzhen Sun , Tianxiang Ma , Songtao Zhao , Siyu Zhou , Qian He

Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffusion models to perform zero-shot point tracking by simply…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Ayush Shrivastava , Sanyam Mehta , Daniel Geng , Andrew Owens

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos,…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Jiayi Gao , Zijin Yin , Changcheng Hua , Yuxin Peng , Kongming Liang , Zhanyu Ma , Jun Guo , Yang Liu

Personalized text-to-image generation has gained significant attention for its capability to generate high-fidelity portraits of specific identities conditioned on user-defined prompts. Existing methods typically involve test-time…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Yujia Wu , Yiming Shi , Jiwei Wei , Chengwei Sun , Yang Yang , Heng Tao Shen

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object motion, often…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Xiaozhe Li , Kai WU , Siyi Yang , YiZhan Qu , Guohua. Zhang , Zhiyu Chen , Jiayao Li , Jiangchuan Mu , Xiaobin Hu , Wen Fang , Mingliang Xiong , Hao Deng , Qingwen Liu , Gang Li , Bin He

With the impressive progress in diffusion-based text-to-image generation, extending such powerful generative ability to text-to-video raises enormous attention. Existing methods either require large-scale text-video pairs and a large number…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Ruiqi Wu , Liangyu Chen , Tong Yang , Chunle Guo , Chongyi Li , Xiangyu Zhang

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes (e.g., appearance,…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Jiwen Yu , Xiaodong Cun , Chenyang Qi , Yong Zhang , Xintao Wang , Ying Shan , Jian Zhang

Incorporating a customized object into image generation presents an attractive feature in text-to-image generation. However, existing optimization-based and encoder-based methods are hindered by drawbacks such as time-consuming…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Ziyang Yuan , Mingdeng Cao , Xintao Wang , Zhongang Qi , Chun Yuan , Ying Shan

We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Wonjoon Jin , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

Personalized image generation, where reference images of one or more subjects are used to generate their image according to a scene description, has gathered significant interest in the community. However, such generated images suffer from…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Parul Gupta , Abhinav Dhall , Thanh-Toan Do

In this paper, we present a diffusion model-based framework for animating people from a single image for a given target 3D motion sequence. Our approach has two core components: a) learning priors about invisible parts of the human body and…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Boyi Li , Junming Chen , Jathushan Rajasegaran , Yossi Gandelsman , Alexei A. Efros , Jitendra Malik

Recently video diffusion models have emerged as expressive generative tools for high-quality video content creation readily available to general users. However, these models often do not offer precise control over camera poses for video…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Dejia Xu , Weili Nie , Chao Liu , Sifei Liu , Jan Kautz , Zhangyang Wang , Arash Vahdat

Recent one-shot video tuning methods, which fine-tune the network on a specific video based on pre-trained text-to-image models (e.g., Stable Diffusion), are popular in the community because of the flexibility. However, these methods often…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Liang Peng , Haoran Cheng , Zheng Yang , Ruisi Zhao , Linxuan Xia , Chaotian Song , Qinglin Lu , Boxi Wu , Wei Liu

Despite tremendous progress in generating high-quality images using diffusion models, synthesizing a sequence of animated frames that are both photorealistic and temporally coherent is still in its infancy. While off-the-shelf billion-scale…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Songwei Ge , Seungjun Nah , Guilin Liu , Tyler Poon , Andrew Tao , Bryan Catanzaro , David Jacobs , Jia-Bin Huang , Ming-Yu Liu , Yogesh Balaji

Editing portrait videos is a challenging task that requires flexible yet precise control over a wide range of modifications, such as appearance changes, expression edits, or the addition of objects. The key difficulty lies in preserving the…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Sagi Polaczek , Or Patashnik , Ali Mahdavi-Amiri , Daniel Cohen-Or

Recent CLIP-guided 3D optimization methods, such as DreamFields and PureCLIPNeRF, have achieved impressive results in zero-shot text-to-3D synthesis. However, due to scratch training and random initialization without prior knowledge, these…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Jiale Xu , Xintao Wang , Weihao Cheng , Yan-Pei Cao , Ying Shan , Xiaohu Qie , Shenghua Gao

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error,…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Koichi Namekata , Sherwin Bahmani , Ziyi Wu , Yash Kant , Igor Gilitschenski , David B. Lindell

Recent text-to-image generation models have demonstrated incredible success in generating images that faithfully follow input prompts. However, the requirement of using words to describe a desired concept provides limited control over the…

计算机视觉与模式识别 · 计算机科学 2024-01-26 Senthil Purushwalkam , Akash Gokul , Shafiq Joty , Nikhil Naik