English
Related papers

Related papers: FiVE: A Fine-grained Video Editing Benchmark for E…

200 papers

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Xiaoyan Cong , Haotian Yang , Angtian Wang , Yizhi Wang , Yiding Yang , Canyu Zhang , Chongyang Ma

Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating…

Graphics · Computer Science 2025-10-07 Nilay Kumar , Priyansh Bhandari , G. Maragatham

Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yiqing Shi , Yiren Song , Mike Zheng Shou

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xingrui Wang , Xin Li , Yaosi Hu , Hanxin Zhu , Chen Hou , Cuiling Lan , Zhibo Chen

The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing benchmarks largely overlooked implausible scenarios and do not measure audio-visual alignment. We…

Multimedia · Computer Science 2026-05-05 Advait Tilak , Jiwon Choi , Nazifa Mouli , Wei Le

Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate…

Image and Video Processing · Electrical Eng. & Systems 2026-05-26 Benjamin Herb , Steve Göring , Alexander Raake , Rakesh Rao Ramachandra Rao

Diffusion models have proven to be highly effective in generating high-quality images. However, adapting large pre-trained diffusion models to new domains remains an open challenge, which is critical for real-world applications. This paper…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Enze Xie , Lewei Yao , Han Shi , Zhili Liu , Daquan Zhou , Zhaoqiang Liu , Jiawei Li , Zhenguo Li

The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyang He , Jie Wang , Jiangning Zhang , Zhucun Xue , Xingyuan Bu , Qiangpeng Yang , Shilei Wen , Lei Xie

Image diffusion models, trained on massive image collections, have emerged as the most versatile image generator model in terms of quality and diversity. They support inverting real images and conditional (e.g., text) generation, making…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Duygu Ceylan , Chun-Hao Paul Huang , Niloy J. Mitra

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression Variational Autoencoder, Video-VAE, is designed for video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Guoqing Ma , Haoyang Huang , Kun Yan , Liangyu Chen , Nan Duan , Shengming Yin , Changyi Wan , Ranchen Ming , Xiaoniu Song , Xing Chen , Yu Zhou , Deshan Sun , Deyu Zhou , Jian Zhou , Kaijun Tan , Kang An , Mei Chen , Wei Ji , Qiling Wu , Wen Sun , Xin Han , Yanan Wei , Zheng Ge , Aojie Li , Bin Wang , Bizhu Huang , Bo Wang , Brian Li , Changxing Miao , Chen Xu , Chenfei Wu , Chenguang Yu , Dapeng Shi , Dingyuan Hu , Enle Liu , Gang Yu , Ge Yang , Guanzhe Huang , Gulin Yan , Haiyang Feng , Hao Nie , Haonan Jia , Hanpeng Hu , Hanqi Chen , Haolong Yan , Heng Wang , Hongcheng Guo , Huilin Xiong , Huixin Xiong , Jiahao Gong , Jianchang Wu , Jiaoren Wu , Jie Wu , Jie Yang , Jiashuai Liu , Jiashuo Li , Jingyang Zhang , Junjing Guo , Junzhe Lin , Kaixiang Li , Lei Liu , Lei Xia , Liang Zhao , Liguo Tan , Liwen Huang , Liying Shi , Ming Li , Mingliang Li , Muhua Cheng , Na Wang , Qiaohui Chen , Qinglin He , Qiuyan Liang , Quan Sun , Ran Sun , Rui Wang , Shaoliang Pang , Shiliang Yang , Sitong Liu , Siqi Liu , Shuli Gao , Tiancheng Cao , Tianyu Wang , Weipeng Ming , Wenqing He , Xu Zhao , Xuelin Zhang , Xianfang Zeng , Xiaojia Liu , Xuan Yang , Yaqi Dai , Yanbo Yu , Yang Li , Yineng Deng , Yingming Wang , Yilei Wang , Yuanwei Lu , Yu Chen , Yu Luo , Yuchu Luo , Yuhe Yin , Yuheng Feng , Yuxiang Yang , Zecheng Tang , Zekai Zhang , Zidong Yang , Binxing Jiao , Jiansheng Chen , Jing Li , Shuchang Zhou , Xiangyu Zhang , Xinhao Zhang , Yibo Zhu , Heung-Yeung Shum , Daxin Jiang

Large-scale text-to-image diffusion models have achieved unprecedented success in image generation and editing. However, extending this success to video editing remains challenging. Recent video editing efforts have adapted pretrained…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Mingshu Cai , Yixuan Li , Osamu Yoshie , Yuya Ieiri

All current benchmarks for multimodal deepfake detection manipulate entire frames using various generation techniques, resulting in oversaturated detection accuracies exceeding 94% at the video-level classification. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Juho Jung , Sangyoun Lee , Jooeon Kang , Yunjin Na

This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Chaeyoung Jung , Suyeon Lee , Ji-Hoon Kim , Joon Son Chung

Monocular Depth Estimation (MDE) is a fundamental 3D vision problem with numerous applications such as 3D scene reconstruction, autonomous navigation, and AI content creation. However, robust and generalizable MDE remains challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Yunpeng Bai , Qixing Huang

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Xuyang Shen , Dong Li , Jinxing Zhou , Zhen Qin , Bowen He , Xiaodong Han , Aixuan Li , Yuchao Dai , Lingpeng Kong , Meng Wang , Yu Qiao , Yiran Zhong

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic…

In this paper, we present CCEdit, a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control, ensuring precise and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Ruoyu Feng , Wenming Weng , Yanhui Wang , Yuhui Yuan , Jianmin Bao , Chong Luo , Zhibo Chen , Baining Guo

Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompts rather than well-designed ones. The primary issue lies in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Xiangjun Zhang , Litong Gong , Yinglin Zheng , Yansong Liu , Wentao Jiang , Mingyi Xu , Biao Wang , Tiezheng Ge , Ming Zeng

This paper proposes the synthetic long-video meta-evaluation (SLVMEval), a benchmark for meta-evaluating text-to-video (T2V) evaluation systems. The proposed SLVMEval benchmark focuses on assessing these systems on videos of up to 10,486 s…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ryosuke Matsuda , Keito Kudo , Haruto Yoshida , Nobuyuki Shimizu , Jun Suzuki

Instruction-based video editing requires transforming a source video according to a natural-language instruction while preserving irrelevant content and remaining temporally coherent. We argue that existing Diffusion Transformer (DiT)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yan Li , Lin Liu , Xiaopeng Zhang , Qi Tian