ASR 系统组合的注意力、CTC、因子化混合与 Transducer 模型比较分析
声音
2025-08-14 v1
摘要
语音识别(ASR)系统的组合方法涵盖了结构化句子级或词级合并技术以及在束搜索期间组合模型得分的方法。本文比较了不同 ASR 架构之间的模型组合。我们的方法利用不同模型在探索搜索空间不同部分方面的互补优势。我们对两个模型候选者的联合假设列表进行重排序。然后通过这些序列级得分的对数线性组合来识别最佳假设。虽然在首次识别期间进行模型组合可能产生改进的性能,但会引入由于解码方法不同导致的变异性,使直接比较更加困难。我们的方法确保了本研究中呈现的所有系统组合结果之间的一致性比较。我们评估了具有不同架构和标签拓扑结构以及单元的模型对候选者。本研究在 Librispeech 960h 任务上提供了实验结果。
关键词
引用
@article{arxiv.2508.09880,
title = {A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models},
author = {Noureldin Bayoumi and Robin Schmitt and Tina Raissi and Albert Zeyer and Ralf Schlüter and Hermann Ney},
journal= {arXiv preprint arXiv:2508.09880},
year = {2025}
}
备注
Accepted for presentation at IEEE Speech Communication; 16th ITG Conference