中文
相关论文

相关论文: Tuning-free Instruction-based Video Editing Via St…

200 篇论文

Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Sunjae Yoon , Gwanhyeong Koo , Ji Woo Hong , Chang D. Yoo

Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approaches rely on fixed Gaussian noise to construct the source trajectory, leading to biased trajectory…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Lifan Jiang , Boxi Wu , Yuhang Pei , Tianrun Wu , Yongyuan Chen , Yan Zhao , Shiyu Yu , Deng Cai

We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $\rho$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Saemee Choi , Sohyun Jeong , Hyojin Jang , Jaegul Choo , Jinhee Kim

Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tuning large-scale video models for specific tasks, which becomes…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Sangwon Jang , Taekyung Ki , Jaehyeong Jo , Jaehong Yoon , Soo Ye Kim , Zhe Lin , Sung Ju Hwang

Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurately, especially when…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Jialu Li , Shoubin Yu , Han Lin , Jaemin Cho , Jaehong Yoon , Mohit Bansal

Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhefan Rao , Bin Zou , Haoxuan Che , Xuanhua He , Chong Hou Choi , Yanheng Li , Rui Liu , Qifeng Chen

With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Haonan Qiu , Menghan Xia , Yong Zhang , Yingqing He , Xintao Wang , Ying Shan , Ziwei Liu

Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Ofir Abramovich , Nadav Z. Cohen , Adi Rosenthal , Ariel Shamir

Despite the rapid progress of instruction-based image editing, its extension to video remains underexplored, primarily due to the prohibitive cost and complexity of constructing large-scale paired video editing datasets. To address this…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Xinyao Liao , Xianfang Zeng , Ziye Song , Zhoujie Fu , Gang Yu , Guosheng Lin

Image denoising is a fundamental task in low-level computer vision. While recent deep learning-based image denoising methods have achieved impressive performance, they are black-box models and the underlying denoising principle remains…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Jingwei Niu , Jun Cheng , Shan Tan

Event-based sensors offer significant advantages over traditional frame-based cameras, especially in scenarios involving rapid motion or challenging lighting conditions. However, event data frequently suffers from considerable noise,…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Marcin Kowalczyk , Kamil Jeziorek , Tomasz Kryjak

With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Kuan-Yu Chen , Jeng-Lin Li , De-Yan Lu , Jian-Jiun Ding

Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Xinyu Chen , Yuyi Qian , Jiang Lin , Shenyi Wang , Gao Wang , Zhiqiu Zhang , Jizhi Zhang , Mingjie Wang , Qiang Tang , Qian Wang , Song Wu , Zili Yi

Though diffusion-based video generation has witnessed rapid progress, the inference results of existing models still exhibit unsatisfactory temporal consistency and unnatural dynamics. In this paper, we delve deep into the noise…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Tianxing Wu , Chenyang Si , Yuming Jiang , Ziqi Huang , Ziwei Liu

Modeling the processing chain that has produced a video is a difficult reverse engineering task, even when the camera is available. This makes model based video processing a still more complex task. In this paper we propose a fully blind…

计算机视觉与模式识别 · 计算机科学 2020-02-26 Thibaud Ehret , Axel Davy , Jean-Michel Morel , Gabriele Facciolo , Pablo Arias

Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of different modalities. In this work, we introduce structured…

机器学习 · 计算机科学 2025-03-21 Aritra Bhowmik , Fida Mohammad Thoker , Carlos Hinojosa , Bernard Ghanem , Cees G. M. Snoek

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim

Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce…

Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, lighting conditions have received limited attention and…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Ryugo Morita , Stanislav Frolov , Brian Bernhard Moser , Ko Watanabe , Riku Takahashi , Andreas Dengel

Text-conditioned image editing has succeeded in various types of editing based on a diffusion framework. Unfortunately, this success did not carry over to a video, which continues to be challenging. Existing video editing systems are still…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Sunjae Yoon , Gwanhyeong Koo , Ji Woo Hong , Chang D. Yoo
‹ 上一页 1 2 3 10 下一页 ›