English
Related papers

Related papers: HunyuanCustom: A Multimodal-Driven Architecture fo…

200 papers

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich…

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Dianbing Xi , Jiepeng Wang , Yuanzhi Liang , Xi Qiu , Jialun Liu , Hao Pan , Yuchi Huo , Rui Wang , Haibin Huang , Chi Zhang , Xuelong Li

Diffusion based video generation has received extensive attention and achieved considerable success within both the academic and industrial communities. However, current efforts are mainly concentrated on single-objective or single-task…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Ludan Ruan , Lei Tian , Chuanwei Huang , Xu Zhang , Xinyan Xiao

Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Zhefan Rao , Bin Zou , Haoxuan Che , Xuanhua He , Chong Hou Choi , Yanheng Li , Rui Liu , Qifeng Chen

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Tao Wu , Yong Zhang , Xintao Wang , Xianpan Zhou , Guangcong Zheng , Zhongang Qi , Ying Shan , Xi Li

Customized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Chi-Pin Huang , Yen-Siang Wu , Hung-Kai Chung , Kai-Po Chang , Fu-En Yang , Yu-Chiang Frank Wang

With the rapid development of AI-generated content (AIGC), video generation has emerged as one of its most dynamic and impactful subfields. In particular, the advancement of video generation foundation models has led to growing demand for…

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hong Chen , Xin Wang , Yipeng Zhang , Yuwei Zhou , Zeyang Zhang , Siao Tang , Wenwu Zhu

Cascaded video super-resolution has emerged as a promising technique for decoupling the computational burden associated with generating high-resolution videos using large foundation models. Existing studies, however, are largely confined to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Shian Du , Menghan Xia , Chang Liu , Quande Liu , Xintao Wang , Pengfei Wan , Xiangyang Ji

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven creation has achieved great progress with the identity…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Ze Ma , Daquan Zhou , Chun-Hsiao Yeh , Xue-She Wang , Xiuyu Li , Huanrui Yang , Zhen Dong , Kurt Keutzer , Jiashi Feng

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled…

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Minjung Shin , Hyunin Cho , Sooyeon Go , Jin-Hwa Kim , Youngjung Uh

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, most approaches rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Wei-Hua Li , Cheng Sun , Chu-Song Chen

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal models lack essential temporally-structured, detailed, and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yuying Ge , Yixiao Ge , Chen Li , Teng Wang , Junfu Pu , Yizhuo Li , Lu Qiu , Jin Ma , Lisheng Duan , Xinyu Zuo , Jinwen Luo , Weibo Gu , Zexuan Li , Xiaojing Zhang , Yangyu Tao , Han Hu , Di Wang , Ying Shan

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Cong Wei , Quande Liu , Zixuan Ye , Qiulin Wang , Xintao Wang , Pengfei Wan , Kun Gai , Wenhu Chen

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face-attribute…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jiazheng Xing , Fei Du , Hangjie Yuan , Pengwei Liu , Hongbin Xu , Hai Ci , Ruigang Niu , Weihua Chen , Fan Wang , Yong Liu

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuanhao Cai , He Zhang , Xi Chen , Jinbo Xing , Yiwei Hu , Yuqian Zhou , Kai Zhang , Zhifei Zhang , Soo Ye Kim , Tianyu Wang , Yulun Zhang , Xiaokang Yang , Zhe Lin , Alan Yuille

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a…

Sound · Computer Science 2026-05-29 Maomao Li , Zhen Li , Kaipeng Zhang , Guosheng Yin , Zhifeng Li , Dong Xu

In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Donghao Zhou , Guisheng Liu , Hao Yang , Jiatong Li , Jingyu Lin , Xiaohu Huang , Yichen Liu , Xin Gao , Cunjian Chen , Shilei Wen , Chi-Wing Fu , Pheng-Ann Heng

Customized generation using diffusion models has made impressive progress in image generation, but remains unsatisfactory in the challenging video generation task, as it requires the controllability of both subjects and motions. To that…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yujie Wei , Shiwei Zhang , Zhiwu Qing , Hangjie Yuan , Zhiheng Liu , Yu Liu , Yingya Zhang , Jingren Zhou , Hongming Shan