English
Related papers

Related papers: EraserDiT: Fast Video Inpainting with Diffusion Tr…

200 papers

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Bhishma Dedhia , David Bourgin , Krishna Kumar Singh , Yuheng Li , Yan Kang , Zhan Xu , Niraj K. Jha , Yuchen Liu

Recent video inpainting methods have achieved encouraging improvements by leveraging optical flow to guide pixel propagation from reference frames either in the image space or feature space. However, they would produce severe artifacts in…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Chaohao Xie , Kai Han , Kwan-Yee K. Wong

Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Qijun Gan , Yi Ren , Chen Zhang , Zhenhui Ye , Pan Xie , Xiang Yin , Zehuan Yuan , Bingyue Peng , Jianke Zhu

Diffusion Transformers (DiT) have shown strong performance in video generation tasks, but their high computational cost makes them impractical for resource-constrained devices like smartphones, and practical on-device generation is even…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Yushu Wu , Yanyu Li , Anil Kag , Ivan Skorokhodov , Willi Menapace , Ke Ma , Arpit Sahni , Ju Hu , Aliaksandr Siarohin , Dhritiman Sagar , Yanzhi Wang , Sergey Tulyakov

Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, extending them to video remains challenging for two reasons:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Jintong Hu , Bin Chen , Zhenyu Hu , Jiayue Liu , Guo Wang , Lu Qi

While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yuxi Liu , Yipeng Hu , Zekun Zhang , Kunze Jiang , Kun Yuan

Video inpainting has been challenged by complex scenarios like large movements and low-light conditions. Current methods, including emerging diffusion models, face limitations in quality and efficiency. This paper introduces the Flow-Guided…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Bohai Gu , Yongsheng Yu , Heng Fan , Libo Zhang

Erase inpainting, or object removal, aims to precisely remove target objects within masked regions while preserving the overall consistency of the surrounding content. Despite diffusion-based methods have made significant strides in the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Yi Liu , Hao Zhou , Wenxiang Shang , Ran Lin , Benlei Cui

Current video deblurring methods have limitations in recovering high-frequency information since the regression losses are conservative with high-frequency details. Since Diffusion Models (DMs) have strong capabilities in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Chen Rao , Guangyuan Li , Zehua Lan , Jiakai Sun , Junsheng Luan , Wei Xing , Lei Zhao , Huaizhong Lin , Jianfeng Dong , Dalong Zhang

Diffusion Transformers (DiT) have revolutionized high-fidelity image and video synthesis, yet their computational demands remain prohibitive for real-time applications. To solve this problem, feature caching has been proposed to accelerate…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Jiacheng Liu , Chang Zou , Yuanhuiyi Lyu , Junjie Chen , Linfeng Zhang

Recent research explores the potential of Diffusion Models (DMs) for consistent object editing, which aims to modify object position, size, and composition, etc., while preserving the consistency of objects and background without changing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Liyao Jiang , Negar Hassanpour , Mohammad Salameh , Mohammadreza Samadi , Jiao He , Fengyu Sun , Di Niu

While diffusion-based image restoration (IR) methods have achieved remarkable success, they are still limited by the low inference speed attributed to the necessity of executing hundreds or even thousands of sampling steps. Existing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Zongsheng Yue , Jianyi Wang , Chen Change Loy

Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Yi Zuo , Lingling Li , Licheng Jiao , Fang Liu , Xu Liu , Wenping Ma , Shuyuan Yang , Yuwei Guo

High-quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views - often due to real-time constraints - leads to missing information and incomplete surfaces in the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Leif Van Holland , Domenic Zingsheim , Mana Takhsha , Hannah Dröge , Patrick Stotko , Markus Plack , Reinhard Klein

We present a novel task called online video editing, which is designed to edit \textbf{streaming} frames while maintaining temporal consistency. Unlike existing offline video editing assuming all frames are pre-established and accessible,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Feng Chen , Zhen Yang , Bohan Zhuang , Qi Wu

Video try-on stands as a promising area for its tremendous real-world potential. Prior works are limited to transferring product clothing images onto person videos with simple poses and backgrounds, while underperforming on casually…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Jun Zheng , Fuwei Zhao , Youjiang Xu , Xin Dong , Xiaodan Liang

Diffusion Transformers (DiT) have emerged as a powerful architecture for image and video generation, offering superior quality and scalability. However, their practical application suffers from inherent dynamic feature instability, leading…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Guanjie Chen , Xinyu Zhao , Yucheng Zhou , Xiaoye Qu , Tianlong Chen , Yu Cheng

Inspired by the impressive performance of recent face image editing methods, several studies have been naturally proposed to extend these methods to the face video editing task. One of the main challenges here is temporal consistency among…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Gyeongman Kim , Hajin Shim , Hyunsu Kim , Yunjey Choi , Junho Kim , Eunho Yang

Diffusion Transformers (DiTs) have achieved state-of-the-art performance in image and video generation, but their success comes at the cost of heavy computation. This inefficiency is largely due to the fixed tokenization process, which uses…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Dahye Kim , Deepti Ghadiyaram , Raghudeep Gadde

Diffusion models have emerged as highly effective techniques for inpainting, however, they remain constrained by slow sampling rates. While recent advances have enhanced generation quality, they have also increased sampling time, thereby…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Tsiry Mayet , Pourya Shamsolmoali , Simon Bernard , Eric Granger , Romain Hérault , Clement Chatelain