English
Related papers

Related papers: MovieFactory: Automatic Movie Creation from Text u…

200 papers

Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. Unlike standard video captioning, it involves not only describing key visual details but also inferring plots that unfold…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Zihao Yue , Yepeng Zhang , Ziheng Wang , Qin Jin

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Ziyang Chen , Prem Seetharaman , Bryan Russell , Oriol Nieto , David Bourgin , Andrew Owens , Justin Salamon

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

Sound · Computer Science 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Junfu Pu , Yuxin Chen , Teng Wang , Ying Shan

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Automatically generating animation from natural language text finds application in a number of areas e.g. movie script writing, instructional videos, and public safety. However, translating natural language text into animation is a…

Computation and Language · Computer Science 2019-04-12 Yeyao Zhang , Eleftheria Tsipidi , Sasha Schriber , Mubbasir Kapadia , Markus Gross , Ashutosh Modi

With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Sibo Zhang , Jiahong Yuan , Miao Liao , Liangjun Zhang

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Bo Peng , Xinyuan Chen , Yaohui Wang , Chaochao Lu , Yu Qiao

We introduce CinemaWorld, a generative augmented reality system that augments the viewer's physical surroundings with automatically generated mixed reality 3D content extracted from and synchronized with 2D movie scenes. Our system…

Human-Computer Interaction · Computer Science 2026-03-10 Keiichi Ihara , DaeHo Lee , Manato Abe , Hye-Young Jo , Ryo Suzuki

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating…

Artificial Intelligence · Computer Science 2026-04-13 Haobo Hu , Qi Mao , Yuanhang Li , Libiao Jin

Advances in technology have led to the development of methods that can create desired visual multimedia. In particular, image generation using deep learning has been extensively studied across diverse fields. In comparison, video…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Doyeon Kim , Donggyu Joo , Junmo Kim

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting sounds that may not…

Sound · Computer Science 2023-11-10 David Chuan-En Lin , Nikolas Martelaro

In this paper, we present VideoGen, a text-to-video generation approach, which can generate a high-definition video with high frame fidelity and strong temporal consistency using reference-guided latent diffusion. We leverage an…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Xin Li , Wenqing Chu , Ye Wu , Weihang Yuan , Fanglong Liu , Qi Zhang , Fu Li , Haocheng Feng , Errui Ding , Jingdong Wang

Text-to-video generation has been dominated by diffusion-based or autoregressive models. These novel models provide plausible versatility, but are criticized for improper physical motion, shading and illumination, camera motion, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Liu He , Yizhi Song , Hejun Huang , Pinxin Liu , Yunlong Tang , Daniel Aliaga , Xin Zhou

We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Rafail Fridman , Amit Abecasis , Yoni Kasten , Tali Dekel

We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Alexey Bokhovkin , Quan Meng , Shubham Tulsiani , Angela Dai

We introduce Text2Cinemagraph, a fully automated method for creating cinemagraphs from text descriptions - an especially challenging task when prompts feature imaginary elements and artistic styles, given the complexity of interpreting the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Aniruddha Mahapatra , Aliaksandr Siarohin , Hsin-Ying Lee , Sergey Tulyakov , Jun-Yan Zhu