English

Listen to Look: Action Recognition by Previewing Audio

Computer Vision and Pattern Recognition 2020-03-31 v3 Machine Learning Sound Audio and Speech Processing

Abstract

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both short-term and long-term visual redundancies. First, we devise an ImgAud2Vid framework that hallucinates clip-level features by distilling from lighter modalities---a single frame and its accompanying audio---reducing short-term temporal redundancy for efficient clip-level recognition. Second, building on ImgAud2Vid, we further propose ImgAud-Skimming, an attention-based long short-term memory network that iteratively selects useful moments in untrimmed videos, reducing long-term temporal redundancy for efficient video-level recognition. Extensive experiments on four action recognition datasets demonstrate that our method achieves the state-of-the-art in terms of both recognition accuracy and speed.

Keywords

Cite

@article{arxiv.1912.04487,
  title  = {Listen to Look: Action Recognition by Previewing Audio},
  author = {Ruohan Gao and Tae-Hyun Oh and Kristen Grauman and Lorenzo Torresani},
  journal= {arXiv preprint arXiv:1912.04487},
  year   = {2020}
}

Comments

Appears in CVPR 2020; Project page: http://vision.cs.utexas.edu/projects/listen_to_look/

R2 v1 2026-06-23T12:40:56.618Z