English
Related papers

Related papers: ModelScope Text-to-Video Technical Report

200 papers

Text-to-Time Series generation holds significant potential to address challenges such as data sparsity, imbalance, and limited availability of multimodal time series datasets across domains. While diffusion models have achieved remarkable…

Machine Learning · Computer Science 2025-05-09 Yunfeng Ge , Jiawei Li , Yiji Zhao , Haomin Wen , Zhao Li , Meikang Qiu , Hongyan Li , Ming Jin , Shirui Pan

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Ernie Chu , Tzuhsuan Huang , Shuo-Yen Lin , Jun-Cheng Chen

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces…

Artificial Intelligence · Computer Science 2025-05-20 Xinlong Chen , Yuanxing Zhang , Chongling Rao , Yushuo Guan , Jiaheng Liu , Fuzheng Zhang , Chengru Song , Qiang Liu , Di Zhang , Tieniu Tan

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zillur Rahman , Alex Sheng , Cristian Meo

Diffusion models have obtained substantial progress in image-to-video generation. However, in this paper, we find that these models tend to generate videos with less motion than expected. We attribute this to the issue called conditional…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Min Zhao , Hongzhou Zhu , Chendong Xiang , Kaiwen Zheng , Chongxuan Li , Jun Zhu

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

Diffusion models developed on top of powerful text-to-image generation models like Stable Diffusion achieve remarkable success in visual story generation. However, the best-performing approach considers historically generated results as…

Computer Vision and Pattern Recognition · Computer Science 2023-05-29 Zhangyin Feng , Yuchen Ren , Xinmiao Yu , Xiaocheng Feng , Duyu Tang , Shuming Shi , Bing Qin

Spatial convolutions are extensively used in numerous deep video models. It fundamentally assumes spatio-temporal invariance, i.e., using shared weights for every location in different frames. This work presents Temporally-Adaptive…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Ziyuan Huang , Shiwei Zhang , Liang Pan , Zhiwu Qing , Yingya Zhang , Ziwei Liu , Marcelo H. Ang

Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurately, especially when…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jialu Li , Shoubin Yu , Han Lin , Jaemin Cho , Jaehong Yoon , Mohit Bansal

Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Xin You , Minghui Zhang , Hanxiao Zhang , Jie Yang , Nassir Navab

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Shijie Wang , Samaneh Azadi , Rohit Girdhar , Saketh Rambhatla , Chen Sun , Xi Yin

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ashkan Taghipour , Morteza Ghahremani , Mohammed Bennamoun , Farid Boussaid , Aref Miri Rekavandi , Zinuo Li , Qiuhong Ke , Hamid Laga

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Zuhao Yang , Jiahui Zhang , Yingchen Yu , Shijian Lu , Song Bai

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose RealCompo, a new…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xinchen Zhang , Ling Yang , Yaqi Cai , Zhaochen Yu , Kai-Ni Wang , Jiake Xie , Ye Tian , Minkai Xu , Yong Tang , Yujiu Yang , Bin Cui

Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompts rather than well-designed ones. The primary issue lies in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Xiangjun Zhang , Litong Gong , Yinglin Zheng , Yansong Liu , Wentao Jiang , Mingyi Xu , Biao Wang , Tiezheng Ge , Ming Zeng

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Guy Yariv , Yuval Kirstain , Amit Zohar , Shelly Sheynin , Yaniv Taigman , Yossi Adi , Sagie Benaim , Adam Polyak

Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Yaosi Hu , Chong Luo , Zhenzhong Chen

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Ting-Chun Wang , Ming-Yu Liu , Jun-Yan Zhu , Guilin Liu , Andrew Tao , Jan Kautz , Bryan Catanzaro