English
Related papers

Related papers: CI-VID: A Coherent Interleaved Text-Video Dataset

200 papers

This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-temporal blocks to ensure consistent frame generation and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Jiuniu Wang , Hangjie Yuan , Dayou Chen , Yingya Zhang , Xiang Wang , Shiwei Zhang

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Qiyao Xue , Xiangyu Yin , Boyuan Yang , Wei Gao

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Text-to-video (T2V) generation has gained significant attention due to its wide applications to video generation, editing, enhancement and translation, \etc. However, high-quality (HQ) video synthesis is extremely challenging because of the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Tao Yang , Yangming Shi , Yunwen Huang , Feng Chen , Yin Zheng , Lei Zhang

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jibin Song , Mingi Kwon , Jaeseok Jeong , Youngjung Uh

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shenghai Yuan , Xianyi He , Yufan Deng , Yang Ye , Jinfa Huang , Bin Lin , Jiebo Luo , Li Yuan

Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) still exhibit…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Jiarui Wang , Huiyu Duan , Ziheng Jia , Yu Zhao , Woo Yi Yang , Zicheng Zhang , Zijian Chen , Juntong Wang , Yuke Xing , Guangtao Zhai , Xiongkuo Min

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ruihao Xi , Xuekuan Wang , Yongcheng Li , Shuhua Li , Zichen Wang , Yiwei Wang , Feng Wei , Cairong Zhao

In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Tongtong Su , Chengyu Wang , Bingyan Liu , Jun Huang , Dongming Lu

Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the state-of-the-art of vid2vid has advanced significantly,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ting-Chun Wang , Ming-Yu Liu , Andrew Tao , Guilin Liu , Jan Kautz , Bryan Catanzaro

Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Luca Zanchetta , Lorenzo Papa , Luca Maiano , Irene Amerini

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yue Ma , Yingqing He , Xiaodong Cun , Xintao Wang , Siran Chen , Ying Shan , Xiu Li , Qifeng Chen

Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Zhao Wang , Aoxue Li , Lingting Zhu , Yong Guo , Qi Dou , Zhenguo Li

Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Ruoxuan Zhang , Jidong Gao , Bin Wen , Hongxia Xie , Chenming Zhang , Hong-Han Shuai , Wen-Huang Cheng

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Hila Chefer , Shiran Zada , Roni Paiss , Ariel Ephrat , Omer Tov , Michael Rubinstein , Lior Wolf , Tali Dekel , Tomer Michaeli , Inbar Mosseri

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., "a woman is drinking water."). Existing TI2V frameworks often require…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Haomiao Ni , Bernhard Egger , Suhas Lohit , Anoop Cherian , Ye Wang , Toshiaki Koike-Akino , Sharon X. Huang , Tim K. Marks

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Tiehan Fan , Kepan Nan , Rui Xie , Penghao Zhou , Zhenheng Yang , Chaoyou Fu , Xiang Li , Jian Yang , Ying Tai

In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Arun Reddy , Alexander Martin , Eugene Yang , Andrew Yates , Kate Sanders , Kenton Murray , Reno Kriz , Celso M. de Melo , Benjamin Van Durme , Rama Chellappa

Learning-based visual data compression and analysis have attracted great interest from both academia and industry recently. More training as well as testing datasets, especially good quality video datasets are highly desirable for related…

Image and Video Processing · Electrical Eng. & Systems 2021-05-14 Xiaozhong Xu , Shan Liu , Zeqiang Li