English

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

Computer Vision and Pattern Recognition 2026-05-11 v2

Abstract

Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks prioritize scale over realistic human depiction. We introduce SynthForensics, a people-centric benchmark of 20,44520{,}445 videos from 8 T2V and 7 I2V open-source generators, paired-source from FF++/DFD reals, two-stage human-validated, in four compression versions with full metadata. In our paired-comparison human study, raters prefer SynthForensics in 7171--77%77\% of head-to-head comparisons against each of nine existing synthetic-video benchmarks, while facial-quality metrics fall within the FF++/DFD baseline range. Across 15 detectors and three protocols, face-based methods drop 1313--5555 AUC points (mean 2727) from FF++ to SynthForensics and a further 2323 under aggressive compression; fine-tuning closes the gap at a backward cost on legacy benchmarks; training from scratch shows synthetic and manipulation features largely disjoint for most detectors. We release dataset, pipeline, and code.

Keywords

Cite

@article{arxiv.2602.04939,
  title  = {SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes},
  author = {Roberto Leotta and Salvatore Alfio Sambataro and Claudio Vittorio Ragaglia and Mirko Casu and Yuri Petralia and Francesco Guarnera and Luca Guarnera and Sebastiano Battiato},
  journal= {arXiv preprint arXiv:2602.04939},
  year   = {2026}
}
R2 v1 2026-07-01T09:36:37.520Z