English

Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol

Computer Vision and Pattern Recognition 2024-05-02 v1 Multimedia Sound Audio and Speech Processing

Abstract

This paper presents a baseline approach and an experimental protocol for a specific content verification problem: detecting discrepancies between the audio and video modalities in multimedia content. We first design and optimize an audio-visual scene classifier, to compare with existing classification baselines that use both modalities. Then, by applying this classifier separately to the audio and the visual modality, we can detect scene-class inconsistencies between them. To facilitate further research and provide a common evaluation platform, we introduce an experimental protocol and a benchmark dataset simulating such inconsistencies. Our approach achieves state-of-the-art results in scene classification and promising outcomes in audio-visual discrepancies detection, highlighting its potential in content verification applications.

Keywords

Cite

@article{arxiv.2405.00384,
  title  = {Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol},
  author = {Konstantinos Apostolidis and Jakob Abesser and Luca Cuccovillo and Vasileios Mezaris},
  journal= {arXiv preprint arXiv:2405.00384},
  year   = {2024}
}

Comments

Accepted for publication, 3rd ACM Int. Workshop on Multimedia AI against Disinformation (MAD'24) at ACM ICMR'24, June 10, 2024, Phuket, Thailand. This is the "accepted version"

R2 v1 2026-06-28T16:12:33.764Z