English

Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains

Sound 2023-07-26 v1 Artificial Intelligence Neural and Evolutionary Computing Audio and Speech Processing Signal Processing

Abstract

Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown VAD and OSD can be trained jointly using a multi-class classification model. However, these works are often restricted to a specific speech domain, lacking information about the generalization capacities of the systems. This paper proposes a complete and new benchmark of different VAD and OSD models, on multiple audio setups (single/multi-channel) and speech domains (e.g. media, meeting...). Our 2/3-class systems, which combine a Temporal Convolutional Network with speech representations adapted to the setup, outperform state-of-the-art results. We show that the joint training of these two tasks offers similar performances in terms of F1-score to two dedicated VAD and OSD systems while reducing the training cost. This unique architecture can also be used for single and multichannel speech processing.

Keywords

Cite

@article{arxiv.2307.13012,
  title  = {Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains},
  author = {Martin Lebourdais and Théo Mariotte and Marie Tahon and Anthony Larcher and Antoine Laurent and Silvio Montresor and Sylvain Meignier and Jean-Hugh Thomas},
  journal= {arXiv preprint arXiv:2307.13012},
  year   = {2023}
}
R2 v1 2026-06-28T11:38:57.403Z