面向低资源设备的端到端非自回归图像到语音系统的压缩
音频与语音处理
2023-12-04 v1 人工智能
计算与语言
计算机视觉与模式识别
图像与视频处理
摘要
视障人士难以使用手机和笔记本等带触摸屏的个人计算设备。图像到语音(ITS)系统可协助他们缓解此问题,但其巨大的模型尺寸使其极难部署于低资源嵌入式设备。本文旨在通过开发高效的端到端神经架构来克服该挑战,以在低资源设备上从显示内容的微小片段生成音频。我们引入了基于 vision transformers 的图像编码器,并利用知识蒸馏将模型从 610 万参数压缩至 246 万参数。人工与自动评估结果均表明,我们的方法仅导致极微小的性能下降,并使推理时间加快 22%。
引用
@article{arxiv.2312.00174,
title = {Compression of end-to-end non-autoregressive image-to-speech system for low-resourced devices},
author = {Gokul Srinivasagan and Michael Deisher and Munir Georges},
journal= {arXiv preprint arXiv:2312.00174},
year = {2023}
}
备注
5 pages, 2 figures, 2 tables, presented at the 15th ITG Conference on Speech Communications, September 2023, Aachen