English
Related papers

Related papers: CVPR 2023 Text Guided Video Editing Competition

200 papers

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Yang Du , Zhuoran Lin , Kaiqiang Song , Biao Wang , Zhicheng Zheng , Tiezheng Ge , Bo Zheng , Qin Jin

Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage video generation as world simulators. Text-to-video (T2V)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Zuhao Liu , Aleksandar Yanev , Ahmad Mahmood , Ivan Nikolov , Saman Motamed , Wei-Shi Zheng , Xi Wang , Lei Sun , Luc Van Gool , Danda Pani Paudel

Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Chi Zhang , Chengjian Feng , Feng Yan , Qiming Zhang , Mingjin Zhang , Yujie Zhong , Jing Zhang , Lin Ma

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Xiaoyan Cong , Haotian Yang , Angtian Wang , Yizhi Wang , Yiding Yang , Canyu Zhang , Chongyang Ma

Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Omkar Thawakar , Dmitry Demidov , Ritesh Thawkar , Rao Muhammad Anwer , Mubarak Shah , Fahad Shahbaz Khan , Salman Khan

Cognitive science has shown that humans perceive videos in terms of events separated by the state changes of dominant subjects. State changes trigger new events and are one of the most useful among the large amount of redundant information…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Yuxuan Wang , Difei Gao , Licheng Yu , Stan Weixian Lei , Matt Feiszli , Mike Zheng Shou

A great video title describes the most salient event compactly and captures the viewer's attention. In contrast, video captioning tends to generate sentences that describe the video as a whole. Although generating a video title…

Computer Vision and Pattern Recognition · Computer Science 2016-09-09 Kuo-Hao Zeng , Tseng-Hung Chen , Juan Carlos Niebles , Min Sun

The rapid advancement of video generation models has made it increasingly challenging to distinguish AI-generated videos from real ones. This issue underscores the urgent need for effective AI-generated video detectors to prevent the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Zhenliang Ni , Qiangyu Yan , Mouxiao Huang , Tianning Yuan , Yehui Tang , Hailin Hu , Xinghao Chen , Yunhe Wang

Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Tsu-Jui Fu , Xin Eric Wang , Scott T. Grafton , Miguel P. Eckstein , William Yang Wang

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ziyang Liu , Kevin Valencia , Justin Cui

Traditional 3D content creation tools empower users to bring their imagination to life by giving them direct control over a scene's geometry, appearance, motion, and camera path. Creating computer-generated videos, however, is a tedious…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shengqu Cai , Duygu Ceylan , Matheus Gadelha , Chun-Hao Paul Huang , Tuanfeng Yang Wang , Gordon Wetzstein

Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question answering is a specific scenario of such AI-human…

Computation and Language · Computer Science 2019-08-01 Guan-Lin Chao , Abhinav Rastogi , Semih Yavuz , Dilek Hakkani-Tür , Jindong Chen , Ian Lane

This paper proposes the synthetic long-video meta-evaluation (SLVMEval), a benchmark for meta-evaluating text-to-video (T2V) evaluation systems. The proposed SLVMEval benchmark focuses on assessing these systems on videos of up to 10,486 s…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ryosuke Matsuda , Keito Kudo , Haruto Yoshida , Nobuyuki Shimizu , Jun Suzuki

This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2025. The aim of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Shuhao Han , Haotian Fan , Fangyuan Kong , Wenjie Liao , Chunle Guo , Chongyi Li , Radu Timofte , Liang Li , Tao Li , Junhui Cui , Yunqiu Wang , Yang Tai , Jingwei Sun , Jianhui Sun , Xinli Yue , Tianyi Wang , Huan Hou , Junda Lu , Xinyang Huang , Zitang Zhou , Zijian Zhang , Xuhui Zheng , Xuecheng Wu , Chong Peng , Xuezhi Cao , Trong-Hieu Nguyen-Mau , Minh-Hoang Le , Minh-Khoa Le-Phan , Duy-Nam Ly , Hai-Dang Nguyen , Minh-Triet Tran , Yukang Lin , Yan Hong , Chuanbiao Song , Siyuan Li , Jun Lan , Zhichao Zhang , Xinyue Li , Wei Sun , Zicheng Zhang , Yunhao Li , Xiaohong Liu , Guangtao Zhai , Zitong Xu , Huiyu Duan , Jiarui Wang , Guangji Ma , Liu Yang , Lu Liu , Qiang Hu , Xiongkuo Min , Zichuan Wang , Zhenchen Tang , Bo Peng , Jing Dong , Fengbin Guan , Zihao Yu , Yiting Lu , Wei Luo , Xin Li , Minhao Lin , Haofeng Chen , Xuanxuan He , Kele Xu , Qisheng Xu , Zijian Gao , Tianjiao Wan , Bo-Cheng Qiu , Chih-Chung Hsu , Chia-ming Lee , Yu-Fan Lin , Bo Yu , Zehao Wang , Da Mu , Mingxiu Chen , Junkang Fang , Huamei Sun , Wending Zhao , Zhiyu Wang , Wang Liu , Weikang Yu , Puhong Duan , Bin Sun , Xudong Kang , Shutao Li , Shuai He , Lingzhi Fu , Heng Cong , Rongyu Zhang , Jiarong He , Zhishan Qiao , Yongqing Huang , Zewen Chen , Zhe Pang , Juan Wang , Jian Guo , Zhizhuo Shao , Ziyu Feng , Bing Li , Weiming Hu , Hesong Li , Dehua Liu , Zeming Liu , Qingsong Xie , Ruichen Wang , Zhihao Li , Yuqi Liang , Jianqi Bi , Jun Luo , Junfeng Yang , Can Li , Jing Fu , Hongwei Xu , Mingrui Long , Lulin Tang

Promotional videos are rapidly becoming a popular medium for persuading people to change their behaviours in many settings (e.g., online shopping, social enterprise initiatives). Today, such videos are often produced by professionals, which…

Multimedia · Computer Science 2021-12-20 Chang Liu , Han Yu

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Generating videos from text has proven to be a significant challenge for existing generative models. We tackle this problem by training a conditional generative model to extract both static and dynamic information from text. This is…

Multimedia · Computer Science 2017-10-03 Yitong Li , Martin Renqiang Min , Dinghan Shen , David Carlson , Lawrence Carin

Recent years have witnessed remarkable progress in 3D content generation. However, corresponding evaluation methods struggle to keep pace. Automatic approaches have proven challenging to align with human preferences, and the mixed…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Weitao Wang , Haoran Xu , Yuxiao Yang , Zhifang Liu , Jun Meng , Haoqian Wang

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity of inter-person interactions. We introduce the task of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yebin Yang , Di Wen , Lei Qi , Weitong Kong , Junwei Zheng , Ruiping Liu , Yufan Chen , Chengzhi Wu , Kailun Yang , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Kunyu Peng

Video editing has garnered increasing attention alongside the rapid progress of diffusion-based video generation models. As part of these advancements, there is a growing demand for more accessible and controllable forms of video editing,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Min-Jung Kim , Dongjin Kim , Seokju Yun , Jaegul Choo
‹ Prev 1 4 5 6 7 8 10 Next ›