English
Related papers

Related papers: MovieFactory: Automatic Movie Creation from Text u…

200 papers

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Leveraging large-scale image-text datasets and advancements in diffusion models, text-driven generative models have made remarkable strides in the field of image generation and editing. This study explores the potential of extending the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Fu-Yun Wang , Wenshuo Chen , Guanglu Song , Han-Jia Ye , Yu Liu , Hongsheng Li

We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Quynh Phung , Long Mai , Fabian David Caba Heilbron , Feng Liu , Jia-Bin Huang , Cusuh Ham

The text-to-video (T2V) generation models, offering convenient visual creation, have recently garnered increasing attention. Despite their substantial potential, the generated videos may present artifacts, including structural…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Jiazi Bu , Pengyang Ling , Pan Zhang , Tong Wu , Xiaoyi Dong , Yuhang Zang , Yuhang Cao , Dahua Lin , Jiaqi Wang

Given a still photograph, one can imagine how dynamic objects might move against a static background. This idea has been actualized in the form of cinemagraphs, where the motion of particular objects within a still image is repeated, giving…

Computer Vision and Pattern Recognition · Computer Science 2018-01-30 Yipin Zhou , Yale Song , Tamara L. Berg

There has been a significant progress in text conditional image generation models. Recent advancements in this field depend not only on improvements in model structures, but also vast quantities of text-image paired datasets. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Seungdae Han , Joohee Kim

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We…

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds…

We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN…

Computer Vision and Pattern Recognition · Computer Science 2022-09-12 Mohammed M. Alghamdi , He Wang , Andrew J. Bulpitt , David C. Hogg

With the advancement of generative models, the synthesis of different sensory elements such as music, visuals, and speech has achieved significant realism. However, the approach to generate multi-sensory outputs has not been fully explored,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Minheng Ni , Chenfei Wu , Huaying Yuan , Zhengyuan Yang , Ming Gong , Lijuan Wang , Zicheng Liu , Wangmeng Zuo , Nan Duan

Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image…

Computer Vision and Pattern Recognition · Computer Science 2016-05-19 Andrew Shin , Katsunori Ohnishi , Tatsuya Harada

Recent video and language pretraining frameworks lack the ability to generate sentences. We present Multimodal Video Generative Pretraining (MV-GPT), a new pretraining framework for learning from unlabelled videos which can be effectively…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Paul Hongsuck Seo , Arsha Nagrani , Anurag Arnab , Cordelia Schmid

In this paper, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes by analyzing the…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Xin Yang , Zongliang Ma , Letian Yu , Ying Cao , Baocai Yin , Xiaopeng Wei , Qiang Zhang , Rynson W. H. Lau

We present Mobius, a novel method to generate seamlessly looping videos from text descriptions directly without any user annotations, thereby creating new visual materials for the multi-media presentation. Our method repurposes the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Xiuli Bi , Jianfei Yuan , Bo Liu , Yong Zhang , Xiaodong Cun , Chi-Man Pun , Bin Xiao

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

Sound · Computer Science 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Taewon Kang , Ming C. Lin

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who"…

Computer Vision and Pattern Recognition · Computer Science 2016-09-22 Makarand Tapaswi , Yukun Zhu , Rainer Stiefelhagen , Antonio Torralba , Raquel Urtasun , Sanja Fidler

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yupeng Zhou , Daquan Zhou , Ming-Ming Cheng , Jiashi Feng , Qibin Hou