English
Related papers

Related papers: Enhancing Video Inpainting with Aligned Frame Inte…

200 papers

We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue. Our core design is a progressive training approach that…

Recent progress in image-to-video (I2V) diffusion models has significantly advanced the field of generative inbetweening, which aims to generate semantically plausible frames between two keyframes. In particular, inference-time sampling…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Wooseok Jeon , Seunghyun Shin , Dongmin Shin , Hae-Gon Jeon

Free-form video inpainting is a very challenging task that could be widely used for video editing such as text removal. Existing patch-based methods could not handle non-repetitive structures such as faces, while directly applying…

Computer Vision and Pattern Recognition · Computer Science 2019-07-24 Ya-Liang Chang , Zhe Yu Liu , Kuan-Ying Lee , Winston Hsu

In the dynamic field of digital content creation using generative models, state-of-the-art video editing models still do not offer the level of quality and control that users desire. Previous works on video editing either extended from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Max Ku , Cong Wei , Weiming Ren , Harry Yang , Wenhu Chen

Video frame interpolation aims to synthesize one or multiple frames between two consecutive frames in a video. It has a wide range of applications including slow-motion video generation, frame-rate up-scaling and developing video codecs.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-14 Saikat Dutta , Arulkumar Subramaniam , Anurag Mittal

High-quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views - often due to real-time constraints - leads to missing information and incomplete surfaces in the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Leif Van Holland , Domenic Zingsheim , Mana Takhsha , Hannah Dröge , Patrick Stotko , Markus Plack , Reinhard Klein

Video diffusion models can generate realistic and temporally consistent videos. This raises concerns about provenance, ownership, and integrity. Watermarking can help address these issues by embedding metadata directly into the content. To…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Mohammadreza Teymoorianfard , Siddarth Sitaraman , Shiqing Ma , Amir Houmansadr

Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Jingxi Chen , Brandon Y. Feng , Haoming Cai , Tianfu Wang , Levi Burner , Dehao Yuan , Cornelia Fermuller , Christopher A. Metzler , Yiannis Aloimonos

Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoupled spatiotemporal process, where the temporal dynamics of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Jiayang Xu , Fan Zhuo , Majun Zhang , Changhao Pan , Zehan Wang , Siyu Chen , Xiaoda Yang , Tao Jin , Zhou Zhao

Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user supervision, for example, when sparse point prompts are provided on a single frame, their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawar Jyoti Deka

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Yue Ma , Yingqing He , Hongfa Wang , Andong Wang , Chenyang Qi , Chengfei Cai , Xiu Li , Zhifeng Li , Heung-Yeung Shum , Wei Liu , Qifeng Chen

In this paper, we present a novel robust framework for low-level vision tasks, including denoising, object removal, frame interpolation, and super-resolution, that does not require any external training data corpus. Our proposed approach…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Gaurav Shrivastava , Ser-Nam Lim , Abhinav Shrivastava

Existing video frame interpolation (VFI) methods often adopt a frame-centric approach, processing videos as independent short segments (e.g., triplets), which leads to temporal inconsistencies and motion artifacts. To overcome this, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Xinyu Peng , Han Li , Yuyang Huang , Ziyang Zheng , Yaoming Wang , Xin Chen , Wenrui Dai , Chenglin Li , Junni Zou , Hongkai Xiong

Predicting future motion trajectories is a critical capability across domains such as robotics, autonomous systems, and human activity forecasting, enabling safer and more intelligent decision-making. This paper proposes a novel, efficient,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Zesen Zhong , Duomin Zhang , Yijia Li

Long-form video editing poses unique challenges due to the exponential increase in the computational cost from joint editing and Denoising Diffusion Implicit Models (DDIM) inversion across extended sequences. To address these limitations,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Mustafa Munir , Md Mostafijur Rahman , Kartikeya Bhardwaj , Paul Whatmough , Radu Marculescu

M3DDM provides a computationally efficient framework for video outpainting via latent diffusion modeling. However, it exhibits significant quality degradation -- manifested as spatial blur and temporal inconsistency -- under challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Takuya Murakawa , Takumi Fukuzawa , Ning Ding , Toru Tamaki

Recently, diffusion models like StableDiffusion have achieved impressive image generation results. However, the generation process of such diffusion models is uncontrollable, which makes it hard to generate videos with continuous and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Zhihao Hu , Dong Xu

Video stabilization refers to the problem of transforming a shaky video into a visually pleasing one. The question of how to strike a good trade-off between visual quality and computational speed has remained one of the open challenges in…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Weiyue Zhao , Xin Li , Zhan Peng , Xianrui Luo , Xinyi Ye , Hao Lu , Zhiguo Cao

Given two consecutive frames, video interpolation aims at generating intermediate frame(s) to form both spatially and temporally coherent video sequences. While most existing methods focus on single-frame interpolation, we propose an…

Computer Vision and Pattern Recognition · Computer Science 2018-07-16 Huaizu Jiang , Deqing Sun , Varun Jampani , Ming-Hsuan Yang , Erik Learned-Miller , Jan Kautz

Recently, memory-based approaches show promising results on semi-supervised video object segmentation. These methods predict object masks frame-by-frame with the help of frequently updated memory of the previous mask. Different from this…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Kwanyong Park , Sanghyun Woo , Seoung Wug Oh , In So Kweon , Joon-Young Lee