English

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

Audio and Speech Processing 2024-05-07 v2 Sound

Abstract

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a TF-GridNet separation architecture, followed by a speaker-agnostic speech recognizer, we achieve state-of-the-art recognition performance in terms of Optimal Reference Combination Word Error Rate (ORC WER). Then, a d-vector-based diarization module is employed to extract speaker embeddings from the enhanced signals and to assign the CSS outputs to the correct speaker. Here, we propose a syntactically informed diarization using sentence- and word-level boundaries of the ASR module to support speaker turn detection. This results in a state-of-the-art Concatenated minimum-Permutation Word Error Rate (cpWER) for the full meeting recognition pipeline.

Keywords

Cite

@article{arxiv.2309.16482,
  title  = {Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization},
  author = {Thilo von Neumann and Christoph Boeddeker and Tobias Cord-Landwehr and Marc Delcroix and Reinhold Haeb-Umbach},
  journal= {arXiv preprint arXiv:2309.16482},
  year   = {2024}
}

Comments

Accepted at HSCMA Sattelite Workshop at ICASSP 2024

R2 v1 2026-06-28T12:34:59.646Z