English
Related papers

Related papers: VideoDirector: Precise Video Editing via Text-to-V…

200 papers

Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Kyungmin Lee , Sangkyung Kwak , Kihyuk Sohn , Jinwoo Shin

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding text-guided video…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Zhixing Zhang , Bichen Wu , Xiaoyan Wang , Yaqiao Luo , Luxin Zhang , Yinan Zhao , Peter Vajda , Dimitris Metaxas , Licheng Yu

Recent Text-to-Video (T2V) models have demonstrated powerful capability in visual simulation of real-world geometry and physical laws, indicating its potential as implicit world models. Inspired by this, we explore the feasibility of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yu Li , Menghan Xia , Gongye Liu , Jianhong Bai , Xintao Wang , Conglang Zhang , Yuxuan Lin , Ruihang Chu , Pengfei Wan , Yujiu Yang

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Youyuan Zhang , Xuan Ju , James J. Clark

Text-guided video prediction (TVP) involves predicting the motion of future frames from the initial frame according to an instruction, which has wide applications in virtual reality, robotics, and content creation. Previous TVP methods make…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Zhen Xing , Qi Dai , Zejia Weng , Zuxuan Wu , Yu-Gang Jiang

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Tao Xia , Jiawei Liu , Yukun Zhang , Ting Liu , Wei Wang , Lei Zhang

Recent advancements in video generation, particularly in diffusion models, have driven notable progress in text-to-video (T2V) and image-to-video (I2V) synthesis. However, challenges remain in effectively integrating dynamic motion signals…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Ziye Li , Hao Luo , Xincheng Shuai , Henghui Ding

Diffusion-based video generation can create realistic videos, yet existing image- and text-based conditioning fails to offer precise motion control. Prior methods for motion-conditioned synthesis typically require model-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Assaf Singer , Noam Rotstein , Amir Mann , Ron Kimmel , Or Litany

Recent text-to-video (T2V) models have demonstrated strong capabilities in producing high-quality, dynamic videos. To improve the visual controllability, recent works have considered fine-tuning pre-trained T2V models to support…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 June Suk Choi , Kyungmin Lee , Sihyun Yu , Yisol Choi , Jinwoo Shin , Kimin Lee

Diffusion-based image-to-video (I2V) models increasingly exhibit world-model-like properties by implicitly capturing temporal dynamics. However, existing studies have mainly focused on visual quality and controllability, and the robustness…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shuhan Xu , Siyuan Liang , Hongling Zheng , Yong Luo , Han Hu , Lefei Zhang , Dacheng Tao

Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term video dependency with self-attention. Unfortunately, they…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Limin Wang , Yu Qiao

Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regions, upgrading amateur footage into professionally styled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhihao Shi , Kejia Yin , Weilin Wan , Yuhongze Zhou , Yuanhao Yu , Xinxin Zuo , Qiang Sun , Juwei Lu

Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions. However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Vatsal Malaviya , Agneet Chatterjee , Maitreya Patel , Yezhou Yang , Chitta Baral

Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existing methods typically add visual conditions to text-to-video (T2V) foundation models by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Bolin Lai , Sangmin Lee , Xu Cao , Xiang Li , James M. Rehg

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Chaohong Guo , Yihan He , Yongwei Nie , Fei Ma , Xuemiao Xu , Chengjiang Long

Large-scale Text-to-Video (T2V) diffusion models have recently demonstrated unprecedented capability to transform natural language descriptions into stunning and photorealistic videos. Despite the promising results, a significant challenge…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Xingyi Yang , Xinchao Wang

Despite significant advancements in video generation and editing using diffusion models, achieving accurate and localized video editing remains a substantial challenge. Additionally, most existing video editing methods primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Chong Mou , Mingdeng Cao , Xintao Wang , Zhaoyang Zhang , Ying Shan , Jian Zhang

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-efficient fine-tuning of image-text pre-trained models, yet they…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Wencheng Zhu , Yuexin Wang , Hongxuan Li , Pengfei Zhu , Qinghua Hu