中文
相关论文

相关论文: MoViE: Mobile Diffusion for Video Editing

200 篇论文

Despite recent progress in diffusion-based video editing, existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Jia-Wei Liu , Yan-Pei Cao , Jay Zhangjie Wu , Weijia Mao , Yuchao Gu , Rui Zhao , Jussi Keppo , Ying Shan , Mike Zheng Shou

With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a risk that malicious individuals might misuse these…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Junjie Cao , Kaizhou Li , Xinchun Yu , Hongxiang Li , Xiaoping Zhang

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdles in (a) accurately…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Hyeonho Jeong , Geon Yeong Park , Jong Chul Ye

Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However, video editing models have not yet reached the same level of visual quality and user control. To address this, we…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Ozgur Kara , Bariscan Kurtkaya , Hidir Yesiltepe , James M. Rehg , Pinar Yanardag

Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task of zero-shot text-to-video generation and propose a low-cost approach (without…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Levon Khachatryan , Andranik Movsisyan , Vahram Tadevosyan , Roberto Henschel , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Video editing is a challenging task that requires manipulating videos on both the spatial and temporal dimensions. Existing methods for video editing mainly focus on changing the appearance or style of the objects in the video, while…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Yao Teng , Enze Xie , Yue Wu , Haoyu Han , Zhenguo Li , Xihui Liu

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparameters. Fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Minghan Li , Chenxi Xie , Yichen Wu , Lei Zhang , Mengyu Wang

Diffusion models are capable of generating impressive images conditioned on text descriptions, and extensions of these models allow users to edit images at a relatively coarse scale. However, the ability to precisely edit the layout,…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Daniel Geng , Andrew Owens

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite extensive efforts, maintaining the temporal consistency of…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Jiangshan Wang , Yue Ma , Jiayi Guo , Yicheng Xiao , Gao Huang , Xiu Li

Generative models, particularly diffusion models, have made significant success in data synthesis across various modalities, including images, videos, and 3D assets. However, current diffusion models are computationally intensive, often…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Yuanzhi Zhu , Hanshu Yan , Huan Yang , Kai Zhang , Junnan Li

The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employing rapid, low-step sampling to generate preliminary outputs…

Recent developments in Video Diffusion Models (VDMs) have demonstrated remarkable capability to generate high-quality video content. Nonetheless, the potential of VDMs for creating transparent videos remains largely uncharted. In this…

图形学 · 计算机科学 2025-03-04 Menghao Li , Zhenghao Zhang , Junchao Liao , Long Qin , Weizhi Wang

Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Trong-Tung Nguyen , Quang Nguyen , Khoi Nguyen , Anh Tran , Cuong Pham

Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding text-guided video…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Zhixing Zhang , Bichen Wu , Xiaoyan Wang , Yaqiao Luo , Luxin Zhang , Yinan Zhao , Peter Vajda , Dimitris Metaxas , Licheng Yu

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Onkar Susladkar , Jishu Sen Gupta , Chirag Sehgal , Sparsh Mittal , Rekha Singhal

While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explored motion controllability as a means to enhance…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Ryan Burgert , Charles Herrmann , Forrester Cole , Michael S Ryoo , Neal Wadhwa , Andrey Voynov , Nataniel Ruiz

Diffusion models have shown remarkable capabilities in generating high-fidelity data across modalities such as images, audio, and video. However, their computational intensity makes deployment on edge devices a significant challenge. This…

分布式、并行与集群计算 · 计算机科学 2025-04-23 Dongqi Zheng

We present a diffusion-based video editing framework, namely DiffusionAtlas, which can achieve both frame consistency and high fidelity in editing video object appearance. Despite the success in image editing, diffusion models still…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Shao-Yu Chang , Hwann-Tzong Chen , Tyng-Luh Liu

This paper addresses the issue of modifying the visual appearance of videos while preserving their motion. A novel framework, named MagicProp, is proposed, which disentangles the video editing process into two stages: appearance editing and…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Hanshu Yan , Jun Hao Liew , Long Mai , Shanchuan Lin , Jiashi Feng

Despite the ability of existing large-scale text-to-image (T2I) models to generate high-quality images from detailed textual descriptions, they often lack the ability to precisely edit the generated or real images. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Chong Mou , Xintao Wang , Jiechong Song , Ying Shan , Jian Zhang