利用上下文数据的两步方法:空中交通通信中的语音识别
摘要
自动语音识别(ASR)作为飞行员与空中交通管制员之间语音通信的辅助手段,可显著降低任务复杂度并提高传输信息的可靠性。ASR 的应用可减少因误解导致的事件数量,并提升空中交通管理(ATM)效率。显然,需要高精度的预测,尤其是关键信息(即呼号与指令)的预测,以最小化错误风险。我们证明,结合 ASR 与自然语言处理(NLP)方法的优势来利用监视数据(即额外模态),有助于显著改善呼号(命名实体)的识别。本文研究一种两步呼号增强方法:(1)在第 1 步(ASR)中,在 G.fst 和/或解码 FST(网格)中降低可能呼号 n-gram 的权重;(2)在第 2 步(NLP)中,利用命名实体识别(NER)从改进后的识别输出中提取呼号,并与监视数据相关联以选择最合适的一个。通过 ASR 与 NLP 方法相结合的呼号 n-gram 增强,最终使呼号识别绝对提升达 53.7%,或相对提升达 60.4%。
引用
@article{arxiv.2202.03725,
title = {A two-step approach to leverage contextual data: speech recognition in air-traffic communications},
author = {Iuliia Nigmatulina and Juan Zuluaga-Gomez and Amrutha Prasad and Seyyed Saeed Sarfjoo and Petr Motlicek},
journal= {arXiv preprint arXiv:2202.03725},
year = {2022}
}
备注
20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. arXiv admin note: text overlap with arXiv:2108.12156