English

Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)

Sound 2019-06-26 v1 Computer Vision and Pattern Recognition Audio and Speech Processing

Abstract

This report describes our submission to the ActivityNet Challenge at CVPR 2019. We use a 3D convolutional neural network (CNN) based front-end and an ensemble of temporal convolution and LSTM classifiers to predict whether a visible person is speaking or not. Our results show significant improvements over the baseline on the AVA-ActiveSpeaker dataset.

Cite

@article{arxiv.1906.10555,
  title  = {Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)},
  author = {Joon Son Chung},
  journal= {arXiv preprint arXiv:1906.10555},
  year   = {2019}
}
R2 v1 2026-06-23T10:03:08.437Z