English
Related papers

Related papers: FLATTEN: optical FLow-guided ATTENtion for consist…

200 papers

We propose a diffusion-based approach for Text-to-Image (T2I) generation with consistent and interactive 3D layout control and editing. While prior methods improve spatial adherence using 2D cues or iterative copy-warp-paste strategies,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Andrea Rigo , Luca Stornaiuolo , Weijie Wang , Mauro Martino , Bruno Lepri , Nicu Sebe

Leveraging pre-trained conditional diffusion models for video editing without further tuning has gained increasing attention due to its promise in film production, advertising, etc. Yet, seminal works in this line fall short in generation…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Zhenyi Liao , Zhijie Deng

Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. However, existing methods often struggle under occlusion,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Lin Liu , Zhihan Xiao , Haohang Xu , Rong Cong , Zhibo Zhang , Xiaopeng Zhang , Qi Tian

Diffusion transformers typically incorporate textual information via attention layers and a modulation mechanism using a pooled text embedding. Nevertheless, recent approaches discard modulation-based text conditioning and rely exclusively…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Nikita Starodubcev , Daniil Pakhomov , Zongze Wu , Ilya Drobyshevskiy , Yuchen Liu , Zhonghao Wang , Yuqian Zhou , Zhe Lin , Dmitry Baranchuk

Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textual prompts; however, while traditional photography offers precise control over camera settings to…

Graphics · Computer Science 2025-06-17 Armando Fortes , Tianyi Wei , Shangchen Zhou , Xingang Pan

Diffusion models have demonstrated their ability to generate diverse and high-quality images, sparking considerable interest in their potential for real image editing applications. However, existing diffusion-based approaches for local…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenjing Huang , Shikui Tu , Lei Xu

Text-guided diffusion models have become essential for high-quality image synthesis, enabling dynamic image editing. In image editing, two crucial aspects are editability, which determines the extent of modification, and faithfulness, which…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Hansam Cho , Seoung Bum Kim

Text-to-video (T2V) synthesis models, such as OpenAI's Sora, have garnered significant attention due to their ability to generate high-quality videos from a text prompt. In diffusion-based T2V models, the attention mechanism is a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Bingyan Liu , Chengyu Wang , Tongtong Su , Huan Ten , Jun Huang , Kailing Guo , Kui Jia

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Yuming Qiao , Fanyi Wang , Jingwen Su , Yanhao Zhang , Yunjie Yu , Siyu Wu , Guo-Jun Qi

Previous text-guided video editing methods often suffer from temporal inconsistency, motion distortion, and-most notably-limited domain transformation. We attribute these limitations to insufficient modeling of spatiotemporal pixel…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Junsung Lee , Junoh Kang , Bohyung Han

Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as they rely on motion representations that are either…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yeji Song , Jaehyun Lee , Mijin Koo , JunHoo Lee , Nojun Kwak

Recent advances in training-free attention control methods have enabled flexible and efficient text-guided editing capabilities for existing generation models. However, current approaches struggle to simultaneously deliver strong editing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Zixin Yin , Ling-Hao Chen , Lionel Ni , Xili Dai

Large-scale text-to-image diffusion models achieve unprecedented success in image generation and editing. However, how to extend such success to video editing is unclear. Recent initial attempts at video editing require significant…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Wen Wang , Yan Jiang , Kangyang Xie , Zide Liu , Hao Chen , Yue Cao , Xinlong Wang , Chunhua Shen

Applying image processing algorithms independently to each frame of a video often leads to undesired inconsistent results over time. Developing temporally consistent video-based extensions, however, requires domain knowledge for individual…

Computer Vision and Pattern Recognition · Computer Science 2018-08-02 Wei-Sheng Lai , Jia-Bin Huang , Oliver Wang , Eli Shechtman , Ersin Yumer , Ming-Hsuan Yang

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving physical consistency…

Recent pre-trained text-to-image flow models have enabled remarkable progress in text-based image editing. Mainstream approaches adopt a corruption-then-restoration paradigm, where the source image is first corrupted into an editable…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yanghao Wang , Zhen Wang , Long Chen

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xingrui Wang , Xin Li , Yaosi Hu , Hanxin Zhu , Chen Hou , Cuiling Lan , Zhibo Chen

Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity, Diffusion Transformers (DiT) remain computationally…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Tianyi Liu , Ye Lu , Linfeng Zhang , Chen Cai , Jianjun Gao , Yi Wang , Kim-Hui Yap , Lap-Pui Chau

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant challenges in handling…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Xiaojing Zhong , Xinyi Huang , Xiaofeng Yang , Guosheng Lin , Qingyao Wu

Despite their generative power, diffusion models struggle to maintain style consistency across images conditioned on the same style prompt, hindering their practical deployment in creative workflows. While several training-free methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Jiexuan Zhang , Yiheng Du , Qian Wang , Weiqi Li , Yu Gu , Jian Zhang
‹ Prev 1 8 9 10 Next ›