English

A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection

Computer Vision and Pattern Recognition 2025-08-29 v1

Abstract

With the rapid advancement of real-time deepfake generation techniques, forged content is becoming increasingly realistic and widespread across applications like video conferencing and social media. Although state-of-the-art detectors achieve high accuracy on standard benchmarks, their heavy computational cost hinders real-time deployment in practical applications. To address this, we propose the Spatial-Frequency Aware Multi-Scale Fusion Network (SFMFNet), a lightweight yet effective architecture for real-time deepfake detection. We design a spatial-frequency hybrid aware module that jointly leverages spatial textures and frequency artifacts through a gated mechanism, enhancing sensitivity to subtle manipulations. A token-selective cross attention mechanism enables efficient multi-level feature interaction, while a residual-enhanced blur pooling structure helps retain key semantic cues during downsampling. Experiments on several benchmark datasets show that SFMFNet achieves a favorable balance between accuracy and efficiency, with strong generalization and practical value for real-time applications.

Keywords

Cite

@article{arxiv.2508.20449,
  title  = {A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection},
  author = {Libo Lv and Tianyi Wang and Mengxiao Huang and Ruixia Liu and Yinglong Wang},
  journal= {arXiv preprint arXiv:2508.20449},
  year   = {2025}
}

Comments

Accepted to PRCV 2025

R2 v1 2026-07-01T05:09:39.586Z