English
Related papers

Related papers: VideoMaker: Zero-shot Customized Video Generation …

200 papers

Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Boyang Wang , Xuweiyi Chen , Matheus Gadelha , Zezhou Cheng

Recent advances in large-scale text-to-image generation models have led to a surge in subject-driven text-to-image generation, which aims to produce customized images that align with textual descriptions while preserving the identity of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kewen Chen , Xiaobin Hu , Wenqi Ren

In the field of 3D content generation, single image scene reconstruction methods still struggle to simultaneously ensure the quality of individual assets and the coherence of the overall scene in complex environments, while texture editing…

Graphics · Computer Science 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance. However, existing methods struggle with a critical trade-off between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zebin Yao , Lei Ren , Huixing Jiang , Wei Chen , Xiaojie Wang , Ruifan Li , Fangxiang Feng

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Aditya Ramesh , Mikhail Pavlov , Gabriel Goh , Scott Gray , Chelsea Voss , Alec Radford , Mark Chen , Ilya Sutskever

Incorporating camera intrinsics into video generation models offers a principled way to control not only scene dynamics but also the imaging process that governs visual appearance. Prior work has primarily focused on extrinsic control, such…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Debabrata Mandal , Zhihan Peng , Yujie Wang , Praneeth Chakravarthula

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Koichi Namekata , Sherwin Bahmani , Ziyi Wu , Yash Kant , Igor Gilitschenski , David B. Lindell

Propagation-based video inpainting using optical flow at the pixel or feature level has recently garnered significant attention. However, it has limitations such as the inaccuracy of optical flow prediction and the propagation of noise over…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Minhyeok Lee , Suhwan Cho , Chajin Shin , Jungho Lee , Sunghun Yang , Sangyoun Lee

Developing effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, producing highly realistic defects remains difficult. We propose…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jaewoo Song , Daemin Park , Kanghyun Baek , Sangyub Lee , Jooyoung Choi , Eunji Kim , Sungroh Yoon

Consistent human-centric image and video synthesis aims to generate images or videos with new poses while preserving appearance consistency with a given reference image, which is crucial for low-cost visual content creation. Recent advances…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Mingdeng Cao , Chong Mou , Ziyang Yuan , Xintao Wang , Zhaoyang Zhang , Ying Shan , Yinqiang Zheng

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Tianqi Zhang , Ziyi Wang , Wenzhao Zheng , Weiliang Chen , Yuanhui Huang , Zhengyang Huang , Jie Zhou , Jiwen Lu

Recent years have witnessed the strong power of 3D generation models, which offer a new level of creative flexibility by allowing users to guide the 3D content generation process through a single image or natural language. However, it…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Fangfu Liu , Hanyang Wang , Weiliang Chen , Haowen Sun , Yueqi Duan

Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yongqi Yang , Huayang Huang , Xu Peng , Xiaobin Hu , Donghao Luo , Jiangning Zhang , Chengjie Wang , Yu Wu

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Yaniv Nikankin , Niv Haim , Michal Irani

Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Edoardo A. Dominici , Thomas Deixelberger , Konstantinos Vardis , Markus Steinberger

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Yuying Ge , Yizhuo Li , Yixiao Ge , Ying Shan

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yuancheng Xu , Wenqi Xian , Li Ma , Julien Philip , Ahmet Levent Taşel , Yiwei Zhao , Ryan Burgert , Mingming He , Oliver Hermann , Oliver Pilarski , Rahul Garg , Paul Debevec , Ning Yu

In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jing He , Haodong Li , Yongzhe Hu , Guibao Shen , Yingjie Cai , Weichao Qiu , Ying-Cong Chen

Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometry, which suffers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jia Li , Han Yan , Yihang Chen , Siqi Li , Xibin Song , Yifu Wang , Jianfei Cai , Tien-Tsin Wong , Pan Ji

This work presents AnyDoor, a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations in a harmonious way. Instead of tuning parameters for each object, our model is trained only…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Xi Chen , Lianghua Huang , Yu Liu , Yujun Shen , Deli Zhao , Hengshuang Zhao
‹ Prev 1 8 9 10 Next ›