English

Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency

Computer Vision and Pattern Recognition 2025-08-04 v1

Abstract

The rapid advancement of diffusion-based video generation models has led to increasingly realistic synthetic content, presenting new challenges for video forgery detection. Existing methods often struggle to capture fine-grained temporal inconsistencies, particularly in AI-generated videos with high visual fidelity and coherent motion. In this work, we propose a detection framework that leverages spatial-temporal consistency by combining RGB appearance features with optical flow residuals. The model adopts a dual-branch architecture, where one branch analyzes RGB frames to detect appearance-level artifacts, while the other processes flow residuals to reveal subtle motion anomalies caused by imperfect temporal synthesis. By integrating these complementary features, the proposed method effectively detects a wide range of forged videos. Extensive experiments on text-to-video and image-to-video tasks across ten diverse generative models demonstrate the robustness and strong generalization ability of the proposed approach.

Keywords

Cite

@article{arxiv.2508.00397,
  title  = {Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency},
  author = {Xi Xue and Kunio Suzuki and Nabarun Goswami and Takuya Shintate},
  journal= {arXiv preprint arXiv:2508.00397},
  year   = {2025}
}
R2 v1 2026-07-01T04:29:01.361Z