Conventional audio-visual models have independent audio and video branches. In this work, we unify the audio and visual branches by designing a Unified Audio-Visual Model (UAVM). The UAVM achieves a new state-of-the-art audio-visual event classification accuracy of 65.8% on VGGSound. More interestingly, we also find a few intriguing properties of UAVM that the modality-independent counterparts do not have.
@article{arxiv.2208.00061,
title = {UAVM: Towards Unifying Audio and Visual Models},
author = {Yuan Gong and Alexander H. Liu and Andrew Rouditchenko and James Glass},
journal= {arXiv preprint arXiv:2208.00061},
year = {2023}
}
Comments
Published in Signal Processing Letters. Code at https://github.com/YuanGongND/uavm