English
Related papers

Related papers: FullDiT: Multi-Task Video Generative Foundation Mo…

200 papers

Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training paradigms that leverage large-scale, unlabeled multimodal data,…

Machine Learning · Computer Science 2025-06-10 Xiaojun Shan , Qi Cao , Xing Han , Haofei Yu , Paul Pu Liang

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual grounding and restricting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yuwei Fang , Willi Menapace , Aliaksandr Siarohin , Tsai-Shien Chen , Kuan-Chien Wang , Ivan Skorokhodov , Graham Neubig , Sergey Tulyakov

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for dynamic interaction,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Yiyuan Zhang , Yuhao Kang , Zhixin Zhang , Xiaohan Ding , Sanyuan Zhao , Xiangyu Yue

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Cong Wei , Quande Liu , Zixuan Ye , Qiulin Wang , Xintao Wang , Pengfei Wan , Kun Gai , Wenhu Chen

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent control. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Yukun Wang , Ruihuang Li , Jiale Tao , Shiyuan Yang , Liyi Chen , Zhantao Yang , Handz , Yulan Guo , Shuai Shao , Qinglin Lu

Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for single-event generation. When handling multi-event prompts,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Qianxun Xu , Chenxi Song , Yujun Cai , Chi Zhang

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yu Lu , Yuanzhi Liang , Linchao Zhu , Yi Yang

Diffusion-based video generation has achieved significant progress, yet generating multiple actions that occur sequentially remains a formidable task. Directly generating a video with sequential actions can be extremely challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Bowen Zhang , Xiaofei Xie , Haotian Lu , Na Ma , Tianlin Li , Qing Guo

Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works often struggle to maintain consistency across multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Xiao Fu , Shitao Tang , Min Shi , Xian Liu , Jinwei Gu , Ming-Yu Liu , Dahua Lin , Chen-Hsuan Lin

Creating large-scale datasets for training high-performance generative models is often prohibitively expensive, especially when associated attributes or annotations must be provided. As a result, merging existing datasets has become a…

Machine Learning · Statistics 2026-03-31 Yanfeng Yang , Kenji Fukumizu

While image manipulation achieves tremendous breakthroughs (e.g., generating realistic faces) in recent years, video generation is much less explored and harder to control, which limits its applications in the real world. For instance,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-08 Tsun-Hsuan Wang , Yen-Chi Cheng , Chieh Hubert Lin , Hwann-Tzong Chen , Min Sun

Instruction-based video editing requires transforming a source video according to a natural-language instruction while preserving irrelevant content and remaining temporally coherent. We argue that existing Diffusion Transformer (DiT)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yan Li , Lin Liu , Xiaopeng Zhang , Qi Tian

Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Zixuan Ye , Xuanhua He , Quande Liu , Qiulin Wang , Xintao Wang , Pengfei Wan , Di Zhang , Kun Gai , Qifeng Chen , Wenhan Luo

High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which is essential for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Yining Yao , Xi Guo , Chenjing Ding , Wei Wu

Due to the statistical complexity of video, the high degree of inherent stochasticity, and the sheer amount of data, generating natural video remains a challenging task. State-of-the-art video generation models often attempt to address…

Computer Vision and Pattern Recognition · Computer Science 2020-02-12 Dirk Weissenborn , Oscar Täckström , Jakob Uszkoreit

Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models often focus on more common categories. In large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Kun Wang , Donglin Di , Tonghua Su , Lei Fan

The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yaofang Liu , Xiaodong Cun , Xuebo Liu , Xintao Wang , Yong Zhang , Haoxin Chen , Yang Liu , Tieyong Zeng , Raymond Chan , Ying Shan

Prompt tuning, which involves training a small set of parameters, effectively enhances the pre-trained Vision-Language Models (VLMs) to downstream tasks. However, they often come at the cost of flexibility and adaptability when the tuned…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Mushui Liu , Bozheng Li , Yunlong Yu

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Yuming Jiang , Tianxing Wu , Shuai Yang , Chenyang Si , Dahua Lin , Yu Qiao , Chen Change Loy , Ziwei Liu

Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Dacheng Li , Yunhao Fang , Yukang Chen , Shuo Yang , Shiyi Cao , Justin Wong , Michael Luo , Xiaolong Wang , Hongxu Yin , Joseph E. Gonzalez , Ion Stoica , Song Han , Yao Lu