English
Related papers

Related papers: Video2LoRA: Unified Semantic-Controlled Video Gene…

200 papers

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xinyang Wang , Qian Liu , Wenjie Ding , Zhao Yang , Wei Li , Chang Liu , Bailin Li , Kun Zhan , Xianpeng Lang , Wei Chen

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Aimon Rahman , Jiang Liu , Ze Wang , Ximeng Sun , Jialian Wu , Xiaodong Yu , Yusheng Su , Vishal M. Patel , Zicheng Liu , Emad Barsoum

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performance degradation due to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Zhen Zhao , Jingqun Tang , Binghong Wu , Chunhui Lin , Shu Wei , Hao Liu , Xin Tan , Zhizhong Zhang , Can Huang , Yuan Xie

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Songcen Xu , Hang Xu , Xiaodan Liang

Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Chenyang Si , Weichen Fan , Zhengyao Lv , Ziqi Huang , Yu Qiao , Ziwei Liu

Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of…

Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a…

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential…

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-range temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Mamshad Nayeem Rizve , Fan Fei , Jayakrishnan Unnikrishnan , Son Tran , Benjamin Z. Yao , Belinda Zeng , Mubarak Shah , Trishul Chilimbi

Low-rank Adaptation (LoRA) has demonstrated remarkable capabilities for task specific fine-tuning. However, in scenarios that involve multiple tasks, training a separate LoRA model for each one results in considerable inefficiency in terms…

Computation and Language · Computer Science 2025-07-08 Yihua Shao , Minxi Yan , Yang Liu , Siyu Chen , Wenjie Chen , Xinwei Long , Ziyang Yan , Lei Li , Chenyu Zhang , Nicu Sebe , Hao Tang , Yan Wang , Hao Zhao , Mengzhu Wang , Jingcai Guo

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xianlong Wang , Wenbo Pan , Shijia Zhou , Ke Li , Yuqi Wang , Zeyu Ye , Hangtao Zhang , Leo Yu Zhang , Xiaohua Jia

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Lei Wang , YuXin Song , Ge Wu , Haocheng Feng , Hang Zhou , Jingdong Wang , Yaxing Wang , jian Yang

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories:…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Hui Han , Siyuan Li , Jiaqi Chen , Yiwen Yuan , Yuling Wu , Chak Tou Leong , Hanwen Du , Junchen Fu , Youhua Li , Jie Zhang , Chi Zhang , Li-jia Li , Yongxin Ni

Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Kaisi Guan , Zhengfeng Lai , Yuchong Sun , Peng Zhang , Wei Liu , Kieran Liu , Meng Cao , Ruihua Song

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Yuchen Hu , Yu Gu , Chenxing Li , Rilin Chen , Dong Yu

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Onkar Susladkar , Jishu Sen Gupta , Chirag Sehgal , Sparsh Mittal , Rekha Singhal

We tackle the long video generation problem, i.e.~generating videos beyond the output length of video generation models. Due to the computation resource constraints, video generation models can only generate video clips that are relatively…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Hsin-Ping Huang , Yu-Chuan Su , Ming-Hsuan Yang