English

Spot the conversation: speaker diarisation in the wild

Sound 2021-08-17 v3 Computer Vision and Pattern Recognition Audio and Speech Processing Image and Video Processing

Abstract

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate videos with diarisation labels. Finally, we use this pipeline to create a large-scale diarisation dataset called VoxConverse, collected from 'in the wild' videos, which we will release publicly to the research community. Our dataset consists of overlapping speech, a large and diverse speaker pool, and challenging background conditions.

Keywords

Cite

@article{arxiv.2007.01216,
  title  = {Spot the conversation: speaker diarisation in the wild},
  author = {Joon Son Chung and Jaesung Huh and Arsha Nagrani and Triantafyllos Afouras and Andrew Zisserman},
  journal= {arXiv preprint arXiv:2007.01216},
  year   = {2021}
}

Comments

The dataset will be available for download from http://www.robots.ox.ac.uk/~vgg/data/voxceleb/voxconverse.html . The development set will be released in July 2020, and the test set will be released in October 2020

R2 v1 2026-06-23T16:48:24.181Z