English

Device Directedness with Contextual Cues for Spoken Dialog Systems

Computation and Language 2022-11-28 v1 Sound Audio and Speech Processing

Abstract

In this work, we define barge-in verification as a supervised learning task where audio-only information is used to classify user spoken dialogue into true and false barge-ins. Following the success of pre-trained models, we use low-level speech representations from a self-supervised representation learning model for our downstream classification task. Further, we propose a novel technique to infuse lexical information directly into speech representations to improve the domain-specific language information implicitly learned during pre-training. Experiments conducted on spoken dialog data show that our proposed model trained to validate barge-in entirely from speech representations is faster by 38% relative and achieves 4.5% relative F1 score improvement over a baseline LSTM model that uses both audio and Automatic Speech Recognition (ASR) 1-best hypotheses. On top of this, our best proposed model with lexically infused representations along with contextual features provides a further relative improvement of 5.7% in the F1 score but only 22% faster than the baseline.

Keywords

Cite

@article{arxiv.2211.13280,
  title  = {Device Directedness with Contextual Cues for Spoken Dialog Systems},
  author = {Dhanush Bekal and Sundararajan Srinivasan and Sravan Bodapati and Srikanth Ronanki and Katrin Kirchhoff},
  journal= {arXiv preprint arXiv:2211.13280},
  year   = {2022}
}
R2 v1 2026-06-28T06:42:52.222Z