English

UAVM: Towards Unifying Audio and Visual Models

Computer Vision and Pattern Recognition 2023-02-17 v2 Multimedia Sound Audio and Speech Processing

Abstract

Conventional audio-visual models have independent audio and video branches. In this work, we unify the audio and visual branches by designing a Unified Audio-Visual Model (UAVM). The UAVM achieves a new state-of-the-art audio-visual event classification accuracy of 65.8% on VGGSound. More interestingly, we also find a few intriguing properties of UAVM that the modality-independent counterparts do not have.

Keywords

Cite

@article{arxiv.2208.00061,
  title  = {UAVM: Towards Unifying Audio and Visual Models},
  author = {Yuan Gong and Alexander H. Liu and Andrew Rouditchenko and James Glass},
  journal= {arXiv preprint arXiv:2208.00061},
  year   = {2023}
}

Comments

Published in Signal Processing Letters. Code at https://github.com/YuanGongND/uavm

R2 v1 2026-06-25T01:20:34.811Z