English
Related papers

Related papers: Tuning-Free Multi-Event Long Video Generation via …

200 papers

Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Penghui Ruan , Pichao Wang , Divya Saxena , Jiannong Cao , Yuhui Shi

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jibin Song , Mingi Kwon , Jaeseok Jeong , Youngjung Uh

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Ryan Burgert , Yuancheng Xu , Wenqi Xian , Oliver Pilarski , Pascal Clausen , Mingming He , Li Ma , Yitong Deng , Lingxiao Li , Mohsen Mousavi , Michael Ryoo , Paul Debevec , Ning Yu

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce MaskFlow, a unified video generation framework that combines…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Michael Fuest , Vincent Tao Hu , Björn Ommer

Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Vasco Ramos , Yonatan Bitton , Michal Yarom , Idan Szpektor , Joao Magalhaes

This paper targets to enhance the diffusion-based text-to-video generation by improving the two input prompts, including the noise and the text. Accommodated with this goal, we propose POS, a training-free Prompt Optimization Suite to boost…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Shijie Ma , Huayi Xu , Mengjian Li , Weidong Geng , Yaxiong Wang , Meng Wang

Generative models have become increasingly powerful tools for robot motion generation, enabling flexible and multimodal trajectory generation across various tasks. Yet, most existing approaches remain limited in handling multiple types of…

Robotics · Computer Science 2026-01-15 Zewen Yang , Xiaobing Dai , Dian Yu , Zhijun Li , Majid Khadiv , Sandra Hirche , Sami Haddadin

Deep generative models provide state-of-the-art performance across a wide array of applications, with recent studies showing increasing applicability for science and engineering. Despite a growing corpus of literature focused on the…

Machine Learning · Computer Science 2026-05-14 Jacob K. Christopher , James E. Warner , Ferdinando Fioretto

Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yongqi Yang , Huayang Huang , Xu Peng , Xiaobin Hu , Donghao Luo , Jiangning Zhang , Chengjie Wang , Yu Wu

Learning to segment images purely by relying on the image-text alignment from web data can lead to sub-optimal performance due to noise in the data. The noise comes from the samples where the associated text does not correlate with the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Yash Patel , Yusheng Xie , Yi Zhu , Srikar Appalaraju , R. Manmatha

Unpaired video-to-video translation aims to translate videos between a source and a target domain without the need of paired training data, making it more feasible for real applications. Unfortunately, the translated videos generally suffer…

Computer Vision and Pattern Recognition · Computer Science 2022-12-22 Kaihong Wang , Kumar Akash , Teruhisa Misu

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yehang Zhang , Xinli Xu , Xiaojie Xu , Li Liu , Yingcong Chen

Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are continuity issues which manifest as object permanence issues, or improperly rendering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Matthew Bendel , Stephen W. Bailey , Mithilesh Vaidya , Sumukh Badam , Xingzhe He

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Shaojin Wu , Mengqi Huang , Wenxu Wu , Yufeng Cheng , Fei Ding , Qian He

Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Thomas Tanay , Mohammed Brahimi , Michal Nazarczuk , Qingwen Zhang , Sibi Catley-Chandar , Arthur Moreau , Zhensong Zhang , Eduardo Pérez-Pellitero

Recently, promptable segmentation models, such as the Segment Anything Model (SAM), have demonstrated robust zero-shot generalization capabilities on static images. These promptable models exhibit denoising abilities for imprecise prompt…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Tao Zhou , Wenhan Luo , Qi Ye , Zhiguo Shi , Jiming Chen

Personalized text-to-image generation aims to create images tailored to user-defined concepts and textual descriptions. Balancing the fidelity of the learned concept with its ability for generation in various contexts presents a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Vera Soboleva , Maksim Nakhodnov , Aibek Alanov

We introduce an audiovisual method for long-range text-to-video retrieval. Unlike previous approaches designed for short video retrieval (e.g., 5-15 seconds in duration), our approach aims to retrieve minute-long videos that capture complex…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Yan-Bo Lin , Jie Lei , Mohit Bansal , Gedas Bertasius

Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is…

Machine Learning · Computer Science 2023-04-05 Jaewoong Lee , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Yunji Kim , Jin-Hwa Kim , Jung-Woo Ha , Sung Ju Hwang