English
Related papers

Related papers: CamI2V: Camera-Controlled Image-to-Video Diffusion…

200 papers

Diffusion models have made significant strides in image generation, mastering tasks such as unconditional image synthesis, text-image translation, and image-to-image conversions. However, their capability falls short in the realm of video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Gaurav Shrivastava , Abhinav Shrivastava

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xinyu Xiao , Binbin Yang , Tingtian Li , Yipeng Yu , Sen Lei

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Min-Jung Kim , Jeongho Kim , Hoiyeong Jin , Junha Hyung , Jaegul Choo

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Yabo Zhang , Yuxiang Wei , Xianhui Lin , Zheng Hui , Peiran Ren , Xuansong Xie , Xiangyang Ji , Wangmeng Zuo

Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Wooseok Jeon , Seungho Park , Seunghyun Shin , Sangeyl Lee , Hyeonho Jeong , Hae-Gon Jeon

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite extensive efforts, maintaining the temporal consistency of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Jiangshan Wang , Yue Ma , Jiayi Guo , Yicheng Xiao , Gao Huang , Xiu Li

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Xiankang He , Guangkai Xu , Bo Zhang , Hao Chen , Ying Cui , Dongyan Guo

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Kanghao Chen , Hangyu Li , JiaZhou Zhou , Zeyu Wang , Lin Wang

Image diffusion models, trained on massive image collections, have emerged as the most versatile image generator model in terms of quality and diversity. They support inverting real images and conditional (e.g., text) generation, making…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Duygu Ceylan , Chun-Hao Paul Huang , Niloy J. Mitra

Recent high-performing image-to-video (I2V) models based on variants of the diffusion transformer (DiT) have displayed remarkable inherent world-modeling capabilities by virtue of training on large scale video datasets. We investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Aaron Appelle , Jerome P. Lynch

Video virtual try-on aims to naturally fit a garment to a target person in consecutive video frames. It is a challenging task, on the one hand, the output video should be in good spatial-temporal consistency, on the other hand, the details…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Cheng Zou , Senlin Cheng , Bolei Xu , Dandan Zheng , Xiaobo Li , Jingdong Chen , Ming Yang

The need for automated real-time visual systems in applications such as smart camera surveillance, smart environments, and drones necessitates the improvement of methods for visual active monitoring and control. Traditionally, the active…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Christos Kyrkou

Event cameras are bio-inspired, motion-activated sensors that demonstrate substantial potential in handling challenging situations, such as motion blur and high-dynamic range. In this paper, we proposed EVI-SAM to tackle the problem of 6…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Weipeng Guan , Peiyu Chen , Huibin Zhao , Yu Wang , Peng Lu

Video generation models have progressed tremendously through large latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Orest Kupyn , Fabian Manhardt , Federico Tombari , Christian Rupprecht

Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camera control in perspective projection video generation, while…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chenhao Ji , Chaohui Yu , Junyao Gao , Fan Wang , Cairong Zhao

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization. We study the design space of data, architecture, and control, and introduce \emph{EasyV2V}, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jinjie Mai , Chaoyang Wang , Guocheng Gordon Qian , Willi Menapace , Sergey Tulyakov , Bernard Ghanem , Peter Wonka , Ashkan Mirzaei

Improving the spatial resolution of CT images is a meaningful yet challenging task, often accompanied by the issue of noise amplification. This article introduces an innovative framework for noise-controlled CT super-resolution utilizing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Yuang Wang , Siyeop Yoon , Rui Hu , Baihui Yu , Duhgoon Lee , Rajiv Gupta , Li Zhang , Zhiqiang Chen , Dufan Wu

Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical…

Camera extrinsic calibration is a fundamental task in computer vision. However, precise relative pose estimation in constrained, highly distorted environments, such as in-cabin automotive monitoring (ICAM), remains challenging. We present…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Felix Stillger , Lukas Hahn , Frederik Hasecke , Tobias Meisen