English

Improving Speech Recognition Accuracy of Local POI Using Geographical Models

Audio and Speech Processing 2021-07-08 v1 Sound

Abstract

Nowadays voice search for points of interest (POI) is becoming increasingly popular. However, speech recognition for local POI has remained to be a challenge due to multi-dialect and massive POI. This paper improves speech recognition accuracy for local POI from two aspects. Firstly, a geographic acoustic model (Geo-AM) is proposed. The Geo-AM deals with multi-dialect problem using dialect-specific input feature and dialect-specific top layer. Secondly, a group of geo-specific language models (Geo-LMs) are integrated into our speech recognition system to improve recognition accuracy of long tail and homophone POI. During decoding, specific language models are selected on demand according to users' geographic location. Experiments show that the proposed Geo-AM achieves 6.5%\sim10.1% relative character error rate (CER) reduction on an accent testset and the proposed Geo-AM and Geo-LM totally achieve over 18.7% relative CER reduction on Tencent Map task.

Keywords

Cite

@article{arxiv.2107.03165,
  title  = {Improving Speech Recognition Accuracy of Local POI Using Geographical Models},
  author = {Songjun Cao and Yike Zhang and Xiaobing Feng and Long Ma},
  journal= {arXiv preprint arXiv:2107.03165},
  year   = {2021}
}

Comments

Accepted by SLT 2021

R2 v1 2026-06-24T03:57:49.083Z