English
Related papers

Related papers: Temporally Consistent Transformers for Video Gener…

200 papers

Real-world videos consist of sequences of events. Generating such sequences with precise temporal control is infeasible with existing video generators that rely on a single paragraph of text as input. When tasked with generating multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ziyi Wu , Aliaksandr Siarohin , Willi Menapace , Ivan Skorokhodov , Yuwei Fang , Varnith Chordia , Igor Gilitschenski , Sergey Tulyakov

Our work explores temporal self-supervision for GAN-based video generation tasks. While adversarial training successfully yields generative models for a variety of areas, temporal relationships in the generated data are much less explored.…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Mengyu Chu , You Xie , Jonas Mayer , Laura Leal-Taixé , Nils Thuerey

Temporal realism remains a central weakness of current generative video models, as most evaluation metrics prioritize spatial appearance and offer limited sensitivity to motion. We introduce a scalable, model-agnostic framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Mert Onur Cakiroglu , Idil Bilge Altun , Zhihe Lu , Mehmet Dalkilic , Hasan Kurban

This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and a short clip that lies inside the long clip. We then extract…

Computer Vision and Pattern Recognition · Computer Science 2021-12-09 Rui Qian , Yeqing Li , Liangzhe Yuan , Boqing Gong , Ting Liu , Matthew Brown , Serge Belongie , Ming-Hsuan Yang , Hartwig Adam , Yin Cui

Realistic temporal dynamics are crucial for many video generation, processing and modelling applications, e.g. in computational fluid dynamics, weather prediction, or long-term climate simulations. Video diffusion models (VDMs) are the…

Machine Learning · Computer Science 2025-05-16 Philipp Hess , Maximilian Gelbrecht , Christof Schötz , Michael Aich , Yu Huang , Shangshang Yang , Niklas Boers

Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans. However, most existing models are trained on short clips, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuchao Gu , Weijia Mao , Mike Zheng Shou

Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity for image…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Xinshuang Liu , Runfa Blark Li , Truong Nguyen

Recently video generation has achieved substantial progress with realistic results. Nevertheless, existing AI-generated videos are usually very short clips ("shot-level") depicting a single scene. To deliver a coherent long video…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Xinyuan Chen , Yaohui Wang , Lingjun Zhang , Shaobin Zhuang , Xin Ma , Jiashuo Yu , Yali Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Face Forgery videos have elicited critical social public concerns and various detectors have been proposed. However, fully-supervised detectors may lead to easily overfitting to specific forgery methods or videos, and existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Daichi Zhang , Zihao Xiao , Shikun Li , Fanzhao Lin , Jianmin Li , Shiming Ge

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds…

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced…

Computer Vision and Pattern Recognition · Computer Science 2018-03-26 Jiawei He , Andreas Lehrmann , Joseph Marino , Greg Mori , Leonid Sigal

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Yanjun Guo , Zhengqiang Zhang , Pengfei Wang , Xinyue Liang , Zhiyuan Ma , Lei Zhang

MoCo is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an input sample, we improve the temporal feature representations…

Computer Vision and Pattern Recognition · Computer Science 2021-03-18 Tian Pan , Yibing Song , Tianyu Yang , Wenhao Jiang , Wei Liu

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. To address this, we propose a training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Shengqu Cai , Weili Nie , Chao Liu , Julius Berner , Lvmin Zhang , Nanye Ma , Hansheng Chen , Maneesh Agrawala , Leonidas Guibas , Gordon Wetzstein , Arash Vahdat

Generating long-duration videos has always been a significant challenge due to the inherent complexity of spatio-temporal domain and the substantial GPU memory demands required to calculate huge size tensors. While diffusion based…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Siyang Zhang , Ser-Nam Lim

Referring video object segmentation aims to segment objects within a video corresponding to a given text description. Existing transformer-based temporal modeling approaches face challenges related to query inconsistency and the limited…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Sun-Hyuk Choi , Hayoung Jo , Seong-Whan Lee

Long video generation remains a challenging and compelling topic in computer vision. Diffusion based models, among the various approaches to video generation, have achieved state of the art quality with their iterative denoising procedures.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Siyang Zhang , Harry Yang , Ser-Nam Lim

Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Harold Haodong Chen , Haojian Huang , Xianfeng Wu , Yexin Liu , Yajing Bai , Wen-Jie Shu , Harry Yang , Ser-Nam Lim

Video temporal grounding aims to pinpoint a video segment that matches the query description. Despite the recent advance in short-form videos (\textit{e.g.}, in minutes), temporal grounding in long videos (\textit{e.g.}, in hours) is still…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yulin Pan , Xiangteng He , Biao Gong , Yiliang Lv , Yujun Shen , Yuxin Peng , Deli Zhao

Video colorization is a challenging and highly ill-posed problem. Although recent years have witnessed remarkable progress in single image colorization, there is relatively less research effort on video colorization and existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Yihao Liu , Hengyuan Zhao , Kelvin C. K. Chan , Xintao Wang , Chen Change Loy , Yu Qiao , Chao Dong