English

Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions

Audio and Speech Processing 2022-02-08 v1 Computation and Language

Abstract

This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) system pipeline. For many low-resource and endangered languages, only single-speaker recordings may be available, demanding a need for domain and speaker-invariant language ID systems. In this memo, we show that a convolutional neural network with a Self-Attentive Pooling layer shows promising results for the language identification task.

Keywords

Cite

@article{arxiv.2104.11985,
  title  = {Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions},
  author = {Roman Bedyakin and Nikolay Mikhaylovskiy},
  journal= {arXiv preprint arXiv:2104.11985},
  year   = {2022}
}

Comments

Accepted to SYGTYP-2021

R2 v1 2026-06-24T01:29:08.945Z