中文
相关论文

相关论文: Mobius: A High Efficient Spatial-Temporal Parallel…

200 篇论文

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Shijie Wang , Samaneh Azadi , Rohit Girdhar , Saketh Rambhatla , Chen Sun , Xi Yin

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Qiyao Xue , Xiangyu Yin , Boyuan Yang , Wei Gao

Generative AI models, particularly Text-to-Video (T2V) systems, offer a promising avenue for transforming science education by automating the creation of engaging and intuitive visual explanations. In this work, we take a first step toward…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Megha Mariam K. M , Aditya Arun , Zakaria Laskar , C. V. Jawahar

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Yang Du , Zhuoran Lin , Kaiqiang Song , Biao Wang , Zhicheng Zheng , Tiezheng Ge , Bo Zheng , Qin Jin

In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained generative T2V model…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Zixin Zhu , Xuelu Feng , Dongdong Chen , Junsong Yuan , Chunming Qiao , Gang Hua

The Text-to-Video (T2V) model aims to generate dynamic and expressive videos from textual prompts. The generation pipeline typically involves multiple modules, such as language encoder, Diffusion Transformer (DiT), and Variational…

分布式、并行与集群计算 · 计算机科学 2025-06-17 Heyang Huang , Cunchen Hu , Jiaqi Zhu , Ziyuan Gao , Liangliang Xu , Yizhou Shan , Yungang Bao , Sun Ninghui , Tianwei Zhang , Sa Wang

This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-temporal blocks to ensure consistent frame generation and…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Jiuniu Wang , Hangjie Yuan , Dayou Chen , Yingya Zhang , Xiang Wang , Shiwei Zhang

Text-to-Time Series generation holds significant potential to address challenges such as data sparsity, imbalance, and limited availability of multimodal time series datasets across domains. While diffusion models have achieved remarkable…

机器学习 · 计算机科学 2025-05-09 Yunfeng Ge , Jiawei Li , Yiji Zhao , Haomin Wen , Zhao Li , Meikang Qiu , Hongyan Li , Ming Jin , Shirui Pan

With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Wenjing Wang , Huan Yang , Zixi Tuo , Huiguo He , Junchen Zhu , Jianlong Fu , Jiaying Liu

State-of-the-art Text-to-Video (T2V) diffusion models can generate visually impressive results, yet they still frequently fail to compose complex scenes or follow logical temporal instructions. In this paper, we argue that many errors,…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Mariam Hassan , Bastien Van Delft , Wuyang Li , Alexandre Alahi

Simultaneous sequence generation is a pivotal task for real-time scenarios, such as streaming speech recognition, simultaneous machine translation and simultaneous speech translation, where the target sequence is generated while receiving…

计算与语言 · 计算机科学 2023-12-01 Shaolei Zhang , Yang Feng

Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xin You , Minghui Zhang , Hanxiao Zhang , Jie Yang , Nassir Navab

In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Jiachen Li , Qian Long , Jian Zheng , Xiaofeng Gao , Robinson Piramuthu , Wenhu Chen , William Yang Wang

Video temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries. Most existing VTG models are built upon frame-wise final-layer CLIP…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Ye Liu , Jixuan He , Wanhua Li , Junsik Kim , Donglai Wei , Hanspeter Pfister , Chang Wen Chen

Huge neural network models have shown unprecedented performance in real-world applications. However, due to memory constraints, model parallelism must be utilized to host large models that would otherwise not fit into the memory of a single…

机器学习 · 计算机科学 2021-04-13 Qifan Xu , Shenggui Li , Chaoyu Gong , Yang You

Most of these text-to-video (T2V) generative models often produce single-scene video clips that depict an entity performing a particular action (e.g., 'a red panda climbing a tree'). However, it is pertinent to generate multi-scene videos…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Hritik Bansal , Yonatan Bitton , Michal Yarom , Idan Szpektor , Aditya Grover , Kai-Wei Chang

The field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalable models. In this work, we present a comprehensive study…

Large diffusion-based Text-to-Image (T2I) models have shown impressive generative powers for text-to-image generation as well as spatially conditioned image generation. For most applications, we can train the model end-toend with paired…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Nithin Gopalakrishnan Nair , Jeya Maria Jose Valanarasu , Vishal M Patel

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zillur Rahman , Alex Sheng , Cristian Meo

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Rui Xie , Yinhong Liu , Penghao Zhou , Chen Zhao , Jun Zhou , Kai Zhang , Zhenyu Zhang , Jian Yang , Zhenheng Yang , Ying Tai