English

MSCT: Differential Cross-Modal Attention for Deepfake Detection

Computer Vision and Pattern Recognition 2026-04-10 v1 Multimedia

Abstract

Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily extract forgery traces through audio-visual alignment, which results from the inconsistency between audio and video modalities. However, the traditional multi-modal forgery detection method has the problem of insufficient feature extraction and modal alignment deviation. To address this, we propose a multi-scale cross-modal transformer encoder (MSCT) for deepfake detection. Our approach includes a multi-scale self-attention to integrate the features of adjacent embeddings and a differential cross-modal attention to fuse multi-modal features. Our experiments demonstrate competitive performance on the FakeAVCeleb dataset, validating the effectiveness of the proposed structure.

Keywords

Cite

@article{arxiv.2604.07741,
  title  = {MSCT: Differential Cross-Modal Attention for Deepfake Detection},
  author = {Fangda Wei and Miao Liu and Yingxue Wang and Jing Wang and Shenghui Zhao and Nan Li},
  journal= {arXiv preprint arXiv:2604.07741},
  year   = {2026}
}

Comments

Accpeted by ICASSP2026

R2 v1 2026-07-01T12:00:26.335Z