English
Related papers

Related papers: CVPR 2023 Text Guided Video Editing Competition

200 papers

The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Yupeng Chen , Penglin Chen , Xiaoyu Zhang , Yixian Huang , Qian Xie

It is still a pipe dream that personal AI assistants on the phone and AR glasses can assist our daily life in addressing our questions like ``how to adjust the date for this watch?'' and ``how to set its heating duration? (while pointing at…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Stan Weixian Lei , Difei Gao , Yuxuan Wang , Dongxing Mao , Zihan Liang , Lingmin Ran , Mike Zheng Shou

Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has received unprecedented attentions over the last few years.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-23 Yitian Yuan , Xiaohan Lan , Xin Wang , Long Chen , Zhi Wang , Wenwu Zhu

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Bosheng Qin , Juncheng Li , Siliang Tang , Tat-Seng Chua , Yueting Zhuang

Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Dawit Mureja Argaw , Fabian Caba Heilbron , Joon-Young Lee , Markus Woodson , In So Kweon

Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text datasets. Recent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Xiefan Guo , Jinlin Liu , Miaomiao Cui , Liefeng Bo , Di Huang

Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Tsu-Jui Fu , Licheng Yu , Ning Zhang , Cheng-Yang Fu , Jong-Chyi Su , William Yang Wang , Sean Bell

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xuan Ju , Tianyu Wang , Yuqian Zhou , He Zhang , Qing Liu , Nanxuan Zhao , Zhifei Zhang , Yijun Li , Yuanhao Cai , Shaoteng Liu , Daniil Pakhomov , Zhe Lin , Soo Ye Kim , Qiang Xu

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employing 2D generation or editing models to process each…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Liyi Chen , Ruihuang Li , Guowen Zhang , Pengfei Wang , Lei Zhang

We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and image inputs. We build Step-Video-TI2V-Eval as a new…

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. However, the open-source…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Keming Wu , Sicong Jiang , Max Ku , Ping Nie , Minghao Liu , Wenhu Chen

In the last few years, we have witnessed a renewed and fast-growing interest in continual learning with deep neural networks with the shared objective of making current AI systems more adaptive, efficient and autonomous. However, despite…

Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Ruotong Liao , Guowen Huang , Qing Cheng , Thomas Seidl , Daniel Cremers , Volker Tresp

The large number of user-generated videos uploaded on to the Internet everyday has led to many commercial video search engines, which mainly rely on text metadata for search. However, metadata is often lacking for user-generated videos,…

Recent text-to-video models have enabled the generation of high-resolution driving scenes from natural language prompts. These AI-generated driving videos (AIGVs) offer a low-cost, scalable alternative to real or simulator data for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Xinhao Xiang , Abhijeet Rastogi , Jiawei Zhang

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to understand video. To…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Hong Zhou , Mike Zheng Shou , Xiang Bai

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiarui Wang , Huiyu Duan , Guangtao Zhai , Juntong Wang , Xiongkuo Min

Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Yuhui Wu , Liyi Chen , Ruibin Li , Shihao Wang , Chenxi Xie , Lei Zhang

With the tremendous growth of videos over the Internet, video thumbnails, providing video content previews, are becoming increasingly crucial to influencing users' online searching experiences. Conventional video thumbnails are generated…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Yitian Yuan , Lin Ma , Wenwu Zhu