English

Development and Evaluation of Video Recordings for the OLSA Matrix Sentence Test

Audio and Speech Processing 2021-04-01 v3 Image and Video Processing

Abstract

One of the established multi-lingual methods for testing speech intelligibility is the matrix sentence test (MST). Most versions of this test are designed with audio-only stimuli. Nevertheless, visual cues play an important role in speech intelligibility, mostly making it easier to understand speech by speechreading. In this work we present the creation and evaluation of dubbed videos for the Oldenburger female MST (OLSA). 28 normal-hearing participants completed test and retest sessions with conditions including audio and visual modalities, speech in quiet and noise, and open and closed-set response formats. The levels to reach 80% sentence intelligibility were measured adaptively for the different conditions. In quiet, the audiovisual benefit compared to audio-only was 7 dB in sound pressure level (SPL). In noise, the audiovisual benefit was 5 dB in signal-to-noise ratio (SNR). Speechreading scores ranged from 0% to 84% speech reception in visual-only sentences, with an average of 50% across participants. This large variability in speechreading abilities was reflected in the audiovisual speech reception thresholds (SRTs), which had a larger standard deviation than the audio-only SRTs. Training and learning effects in audiovisual sentences were found: participants improved their SRTs by approximately 3 dB SNR after 5 trials. Participants retained their best scores on a separate retest session and further improved their SRTs by approx. -1.5 dB.

Keywords

Cite

@article{arxiv.1912.04700,
  title  = {Development and Evaluation of Video Recordings for the OLSA Matrix Sentence Test},
  author = {Gerard Llorach and Frederike Kirschner and Giso Grimm and Melanie A. Zokoll and Kirsten C. Wagener and Volker Hohmann},
  journal= {arXiv preprint arXiv:1912.04700},
  year   = {2021}
}

Comments

10 pages, 9 figures