English

Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision

Sound 2020-11-05 v2 Computer Vision and Pattern Recognition Audio and Speech Processing

Abstract

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal synchrony. We build on earlier work to train embeddings that are more discriminative for uni-modal downstream tasks. To this end, we propose a novel training strategy that not only optimises metrics across modalities, but also enforces intra-class feature separation within each of the modalities. The effectiveness of the method is demonstrated on two downstream tasks: lip reading using the features trained on audio-visual synchronisation, and speaker recognition using the features trained for cross-modal biometric matching. The proposed method outperforms state-of-the-art self-supervised baselines by a signficant margin.

Keywords

Cite

@article{arxiv.2004.14326,
  title  = {Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision},
  author = {Soo-Whan Chung and Hong Goo Kang and Joon Son Chung},
  journal= {arXiv preprint arXiv:2004.14326},
  year   = {2020}
}

Comments

Under submission as a conference paper