English
Related papers

Related papers: DreamVideo-2: Zero-Shot Subject-Driven Video Custo…

200 papers

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Hila Chefer , Shiran Zada , Roni Paiss , Ariel Ephrat , Omer Tov , Michael Rubinstein , Lior Wolf , Tali Dekel , Tomer Michaeli , Inbar Mosseri

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Ruicheng Zhang , Jun Zhou , Zunnan Xu , Zihao Liu , Jiehui Huang , Mingyang Zhang , Yu Sun , Xiu Li

Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately modeling coherent motions…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Yuanpeng Tu , Hao Luo , Xi Chen , Sihui Ji , Xiang Bai , Hengshuang Zhao

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xiaoyu Shi , Zhaoyang Huang , Fu-Yun Wang , Weikang Bian , Dasong Li , Yi Zhang , Manyuan Zhang , Ka Chun Cheung , Simon See , Hongwei Qin , Jifeng Dai , Hongsheng Li

Recent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Xierui Wang , Siming Fu , Qihan Huang , Wanggui He , Hao Jiang

Text-to-video models have demonstrated impressive capabilities in producing diverse and captivating video content, showcasing a notable advancement in generative AI. However, these models generally lack fine-grained control over motion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tuna Han Salih Meral , Hidir Yesiltepe , Connor Dunlop , Pinar Yanardag

Online Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based on the Tracking-by-Detection (TBD) paradigm combined with…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Zhen He , Jian Li , Daxue Liu , Hangen He , David Barber

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Qiang Zhou , Shaofeng Zhang , Nianzu Yang , Ye Qian , Hao Li

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Shuai Yang , Yifan Zhou , Ziwei Liu , Chen Change Loy

We present DiffIR2VR-Zero, a zero-shot framework that enables any pre-trained image restoration diffusion model to perform high-quality video restoration without additional training. While image diffusion models have shown remarkable…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Chang-Han Yeh , Hau-Shiang Shiu , Chin-Yang Lin , Zhixiang Wang , Chi-Wei Hsiao , Ting-Hsuan Chen , Yu-Lun Liu

Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now…

In this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object representation. An…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Tianfei Zhou , Shunzhou Wang , Yi Zhou , Yazhou Yao , Jianwu Li , Ling Shao

Incorporating a customized object into image generation presents an attractive feature in text-to-image generation. However, existing optimization-based and encoder-based methods are hindered by drawbacks such as time-consuming…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Ziyang Yuan , Mingdeng Cao , Xintao Wang , Zhongang Qi , Chun Yuan , Ying Shan

The quadratic time and memory complexity of the attention mechanism in modern Transformer based video generators makes end-to-end training for ultra high resolution videos prohibitively expensive. Motivated by this limitation, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yunfeng Wu , Jiayi Song , Zhenxiong Tan , Zihao He , Songhua Liu

Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subject, or support only…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tsai-Shien Chen , Aliaksandr Siarohin , Willi Menapace , Yuwei Fang , Kwot Sin Lee , Ivan Skorokhodov , Kfir Aberman , Jun-Yan Zhu , Ming-Hsuan Yang , Sergey Tulyakov

Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Zijian Zhou , Shikun Liu , Haozhe Liu , Haonan Qiu , Zhaochong An , Weiming Ren , Zhiheng Liu , Xiaoke Huang , Kam Woh Ng , Tian Xie , Xiao Han , Yuren Cong , Hang Li , Chuyan Zhu , Aditya Patel , Tao Xiang , Sen He

Animation techniques bring digital 3D worlds and characters to life. However, manual animation is tedious and automated techniques are often specialized to narrow shape classes. In our work, we propose a technique for automatic re-animation…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Lukas Uzolas , Elmar Eisemann , Petr Kellnhofer

Incorporating a temporal dimension into pretrained image diffusion models for video generation is a prevalent approach. However, this method is computationally demanding and necessitates large-scale video datasets. More critically, the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Dengsheng Chen , Jie Hu , Xiaoming Wei , Enhua Wu

The Segment Anything Model 2 (SAM 2) has demonstrated strong performance in object segmentation tasks but faces challenges in visual object tracking, particularly when managing crowded scenes with fast-moving or self-occluding objects.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Cheng-Yen Yang , Hsiang-Wei Huang , Wenhao Chai , Zhongyu Jiang , Jenq-Neng Hwang

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents, controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yuxin Zhang , Fan Tang , Nisha Huang , Haibin Huang , Chongyang Ma , Weiming Dong , Changsheng Xu