中文

面向低资源设备的端到端非自回归图像到语音系统的压缩

音频与语音处理 2023-12-04 v1 人工智能 计算与语言 计算机视觉与模式识别 图像与视频处理

摘要

视障人士难以使用手机和笔记本等带触摸屏的个人计算设备。图像到语音(ITS)系统可协助他们缓解此问题,但其巨大的模型尺寸使其极难部署于低资源嵌入式设备。本文旨在通过开发高效的端到端神经架构来克服该挑战,以在低资源设备上从显示内容的微小片段生成音频。我们引入了基于 vision transformers 的图像编码器,并利用知识蒸馏将模型从 610 万参数压缩至 246 万参数。人工与自动评估结果均表明,我们的方法仅导致极微小的性能下降,并使推理时间加快 22%。

关键词

引用

@article{arxiv.2312.00174,
  title  = {Compression of end-to-end non-autoregressive image-to-speech system for low-resourced devices},
  author = {Gokul Srinivasagan and Michael Deisher and Munir Georges},
  journal= {arXiv preprint arXiv:2312.00174},
  year   = {2023}
}

备注

5 pages, 2 figures, 2 tables, presented at the 15th ITG Conference on Speech Communications, September 2023, Aachen