English

OxfordVGG Submission to the EGO4D AV Transcription Challenge

Sound 2023-07-19 v1 Machine Learning Audio and Speech Processing

Abstract

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcription of long-form audio with word-level time alignment, along with two text normalisers which are publicly available. Our final submission obtained 56.0% of the Word Error Rate (WER) on the challenge test set, ranked 1st on the leaderboard. All baseline codes and models are available on https://github.com/m-bain/whisperX.

Keywords

Cite

@article{arxiv.2307.09006,
  title  = {OxfordVGG Submission to the EGO4D AV Transcription Challenge},
  author = {Jaesung Huh and Max Bain and Andrew Zisserman},
  journal= {arXiv preprint arXiv:2307.09006},
  year   = {2023}
}

Comments

Technical Report