English
Related papers

Related papers: Vidu: a Highly Consistent, Dynamic and Skilled Tex…

200 papers

Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Luca Zanchetta , Lorenzo Papa , Luca Maiano , Irene Amerini

We present On-device Sora, the first model training-free solution for diffusion-based on-device text-to-video generation that operates efficiently on smartphone-grade devices. To address the challenges of diffusion-based text-to-video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Bosung Kim , Kyuhwan Lee , Isu Jeong , Jungmin Cheon , Yeojin Lee , Seulki Lee

We present On-device Sora, the first model training-free solution for diffusion-based on-device text-to-video generation that operates efficiently on smartphone-grade devices. To address the challenges of diffusion-based text-to-video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Bosung Kim , Kyuhwan Lee , Isu Jeong , Jungmin Cheon , Yeojin Lee , Seulki Lee

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Canxuan Gang

While recent years have witnessed great progress on using diffusion models for video generation, most of them are simple extensions of image generation frameworks, which fail to explicitly consider one of the key differences between videos…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Jingyun Liang , Yuchen Fan , Kai Zhang , Radu Timofte , Luc Van Gool , Rakesh Ranjan

We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Wonjoon Jin , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Yiming Xie , Chun-Han Yao , Vikram Voleti , Huaizu Jiang , Varun Jampani

Text-to-video generation enhances content creation but is highly computationally intensive: The computational cost of Diffusion Transformers (DiTs) scales quadratically in the number of pixels. This makes minute-length video generation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hongjie Wang , Chih-Yao Ma , Yen-Cheng Liu , Ji Hou , Tao Xu , Jialiang Wang , Felix Juefei-Xu , Yaqiao Luo , Peizhao Zhang , Tingbo Hou , Peter Vajda , Niraj K. Jha , Xiaoliang Dai

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Zirui Pan , Xin Wang , Yipeng Zhang , Hong Chen , Kwan Man Cheng , Yaofei Wu , Wenwu Zhu

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and…

We tackle the long video generation problem, i.e.~generating videos beyond the output length of video generation models. Due to the computation resource constraints, video generation models can only generate video clips that are relatively…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Hsin-Ping Huang , Yu-Chuan Su , Ming-Hsuan Yang

We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Lijun Yu , Yong Cheng , Kihyuk Sohn , José Lezama , Han Zhang , Huiwen Chang , Alexander G. Hauptmann , Ming-Hsuan Yang , Yuan Hao , Irfan Essa , Lu Jiang

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenqi Ouyang , Zeqi Xiao , Danni Yang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, along with other text-to-video diffusion models, is highly reliant on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Wenhao Wang , Yi Yang

AI-generated content has attracted lots of attention recently, but photo-realistic video synthesis is still challenging. Although many attempts using GANs and autoregressive models have been made in this area, the visual quality and length…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Yingqing He , Tianyu Yang , Yong Zhang , Ying Shan , Qifeng Chen

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple image or video diffusion models, utilizing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Hanwen Liang , Yuyang Yin , Dejia Xu , Hanxue Liang , Zhangyang Wang , Konstantinos N. Plataniotis , Yao Zhao , Yunchao Wei

The rapid development of generative models has significantly advanced image and video applications. Among these, video creation, aimed at generating videos under various conditions, has gained substantial attention. However, existing video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Yutong Wang , Haiyu Zhang , Tianfan Xue , Yu Qiao , Yaohui Wang , Chang Xu , Xinyuan Chen

Text-to-video generation has been dominated by diffusion-based or autoregressive models. These novel models provide plausible versatility, but are criticized for improper physical motion, shading and illumination, camera motion, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Liu He , Yizhi Song , Hejun Huang , Pinxin Liu , Yunlong Tang , Daniel Aliaga , Xin Zhou

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ziyang Liu , Kevin Valencia , Justin Cui

While one-step diffusion models have recently excelled in perceptual image compression, their application to video remains limited. Prior efforts typically rely on pretrained 2D autoencoders that generate per-frame latent representations…

Image and Video Processing · Electrical Eng. & Systems 2026-01-06 Xingchen Li , Junzhe Zhang , Junqi Shi , Ming Lu , Zhan Ma