English
Related papers

Related papers: ANYPORTAL: Zero-Shot Consistent Video Background R…

200 papers

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera pose annotations, which are both data-intensive and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zhenghong Zhou , Jie An , Jiebo Luo

Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task of zero-shot text-to-video generation and propose a low-cost approach (without…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Levon Khachatryan , Andranik Movsisyan , Vahram Tadevosyan , Roberto Henschel , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

We present FreeMorph, the first tuning-free method for image morphing that accommodates inputs with different semantics or layouts. Unlike existing methods that rely on finetuning pre-trained diffusion models and are limited by time…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Yukang Cao , Chenyang Si , Jinghao Wang , Ziwei Liu

In the field of 3D content generation, single image scene reconstruction methods still struggle to simultaneously ensure the quality of individual assets and the coherence of the overall scene in complex environments, while texture editing…

Graphics · Computer Science 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Applying an image processing algorithm independently to each video frame often leads to temporal inconsistency in the resulting video. To address this issue, we present a novel and general approach for blind video temporal consistency. Our…

Computer Vision and Pattern Recognition · Computer Science 2022-01-28 Chenyang Lei , Yazhou Xing , Hao Ouyang , Qifeng Chen

Image deocclusion (or amodal completion) aims to recover the invisible regions (\ie, shape and appearance) of occluded instances in images. Despite recent advances, the scarcity of high-quality data that balances diversity, plausibility,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Xinyang Li , Chengjie Yi , Jiawei Lai , Mingbao Lin , Yansong Qu , Shengchuan Zhang , Liujuan Cao

Portrait Animation aims to synthesize a lifelike video from a single source image, using it as an appearance reference, with motion (i.e., facial expressions and head pose) derived from a driving video, audio, text, or generation. Instead…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Jianzhu Guo , Dingyun Zhang , Xiaoqiang Liu , Zhizhou Zhong , Yuan Zhang , Pengfei Wan , Di Zhang

In this paper, we propose a novel framework for solving high-definition video inverse problems using latent image diffusion models. Building on recent advancements in spatio-temporal optimization for video inverse problems using image…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Taesung Kwon , Jong Chul Ye

Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jan Ackermann , Shengqu Cai , Boyang Deng , Zhengfei Kuang , Songyou Peng , Gordon Wetzstein

Benchmarking is central to optimization research, yet existing test suites for continuous optimization remain limited: classical collections are fixed and rigid, while previous generators cover only narrow families of landscapes with…

Neural and Evolutionary Computing · Computer Science 2025-12-02 Danial Yazdani , Mai Peng , Delaram Yazdani , Shima F. Yazdi , Mohammad Nabi Omidvar , Yuan Sun , Trung Thanh Nguyen , Changhe Li , Xiaodong Li

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Junying Wang , Jingyuan Liu , Xin Sun , Krishna Kumar Singh , Zhixin Shu , He Zhang , Jimei Yang , Nanxuan Zhao , Tuanfeng Y. Wang , Simon S. Chen , Ulrich Neumann , Jae Shin Yoon

Diffusion models have revolutionized image generation and editing, producing state-of-the-art results in conditioned and unconditioned image synthesis. While current techniques enable user control over the degree of change in an image edit,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-01 Eran Levin , Ohad Fried

Event cameras are a new type of vision sensor that incorporates asynchronous and independent pixels, offering advantages over traditional frame-based cameras such as high dynamic range and minimal motion blur. However, their output is not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Burak Ercan , Onur Eker , Aykut Erdem , Erkut Erdem

Recent advancements in image relighting models, driven by large-scale datasets and pre-trained diffusion models, have enabled the imposition of consistent lighting. However, video relighting still lags, primarily due to the excessive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Yujie Zhou , Jiazi Bu , Pengyang Ling , Pan Zhang , Tong Wu , Qidong Huang , Jinsong Li , Xiaoyi Dong , Yuhang Zang , Yuhang Cao , Anyi Rao , Jiaqi Wang , Li Niu

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Michal Geyer , Omer Bar-Tal , Shai Bagon , Tali Dekel

Recent advances in video insertion based on diffusion models are impressive. However, existing methods rely on complex control signals but struggle with subject consistency, limiting their practical applicability. In this paper, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Jinshu Chen , Xinghui Li , Xu Bai , Tianxiang Ma , Pengze Zhang , Zhuowei Chen , Gen Li , Lijie Liu , Songtao Zhao , Bingchuan Li , Qian He

Zero-shot depth estimation (DE) models exhibit strong generalization performance as they are trained on large-scale datasets. However, existing models struggle with high-resolution images due to the discrepancy in image resolutions of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Byeongjun Kwon , Munchurl Kim

Following the advancements in text-guided image generation technology exemplified by Stable Diffusion, video generation is gaining increased attention in the academic community. However, relying solely on text guidance for video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Haoyu Zhao , Yuanfan Guo , Jianhua Han , Hang Xu , Xiaodan Liang

We present Depth Anything at Any Condition (DepthAnything-AC), a foundation monocular depth estimation (MDE) model capable of handling diverse environmental conditions. Previous foundation MDE models achieve impressive performance across…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Boyuan Sun , Modi Jin , Bowen Yin , Qibin Hou

Video matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Geng Lin , Chen Gao , Jia-Bin Huang , Changil Kim , Yipeng Wang , Matthias Zwicker , Ayush Saraf