English
Related papers

Related papers: CustomCrafter: Customized Video Generation with Pr…

200 papers

Pre-trained large text-to-image (T2I) models with an appropriate text prompt has attracted growing interests in customized images generation field. However, catastrophic forgetting issue make it hard to continually synthesize new…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Chenxi Liu , Gan Sun , Wenqi Liang , Jiahua Dong , Can Qin , Yang Cong

Remarkable progress has been achieved in image generation with the introduction of generative models. However, precisely controlling the content in generated images remains a challenging task due to their fundamental training objective.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Giang H. Le , Anh Q. Nguyen , Byeongkeun Kang , Yeejin Lee

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Haotian Dong , Xin Wang , Di Lin , Yipeng Wu , Qin Chen , Ruonan Liu , Kairui Yang , Ping Li , Qing Guo

Text-guided generative diffusion models unlock powerful image creation and editing tools. While these have been extended to video generation, current approaches that edit the content of existing footage while retaining structure require…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Patrick Esser , Johnathan Chiu , Parmida Atighehchian , Jonathan Granskog , Anastasis Germanidis

Video data is more cost-effective than motion capture data for learning 3D character motion controllers, yet synthesizing realistic and diverse behaviors directly from videos remains challenging. Previous approaches typically rely on…

Graphics · Computer Science 2025-12-10 Jianan Li , Xiao Chen , Tao Huang , Tien-Tsin Wong

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Yuxuan Bian , Ailing Zeng , Xuan Ju , Xian Liu , Zhaoyang Zhang , Wei Liu , Qiang Xu

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Recent approaches to controllable 4D video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This dominant paradigm is computationally expensive, requiring large-scale datasets and architectural modifications,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yeobin Hong , Suhyeon Lee , Hyungjin Chung , Jong Chul Ye

We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Jinyoung Jun , Won-Dong Jang , Wenbin Ouyang , Raghudeep Gadde , Jungbeom Lee

Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Zeqi Xiao , Yifan Zhou , Shuai Yang , Xingang Pan

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal…

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao

The customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Zebin Yao , Fangxiang Feng , Ruifan Li , Xiaojie Wang

Current advances in human head modeling allow the generation of plausible-looking 3D head models via neural representations, such as NeRFs and SDFs. Nevertheless, constructing complete high-fidelity head models with explicitly controlled…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Artem Sevastopolsky , Philip-William Grassal , Simon Giebenhain , ShahRukh Athar , Luisa Verdoliva , Matthias Niessner

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yifang Men , Yuan Yao , Miaomiao Cui , Liefeng Bo

Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-refinement framework…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yang Liu , Xilin Zhao , Peisong Wen , Siran Dai , Qingming Huang

Consistent human-centric image and video synthesis aims to generate images or videos with new poses while preserving appearance consistency with a given reference image, which is crucial for low-cost visual content creation. Recent advances…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Mingdeng Cao , Chong Mou , Ziyang Yuan , Xintao Wang , Zhaoyang Zhang , Ying Shan , Yinqiang Zheng

Recent advancements in personalizing text-to-image (T2I) diffusion models have shown the capability to generate images based on personalized visual concepts using a limited number of user-provided examples. However, these models often…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yan Hong , Jianfu Zhang

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance. However, existing methods struggle with a critical trade-off between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zebin Yao , Lei Ren , Huixing Jiang , Wei Chen , Xiaojie Wang , Ruifan Li , Fangxiang Feng

Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Dong-Yu Chen , Yixin Guo , Shuojin Yang , Tai-Jiang Mu , Shi-Min Hu
‹ Prev 1 4 5 6 7 8 10 Next ›