面向实现人类水平的端到端同时语音翻译
计算与语言
2024-09-02 v2 声音
音频与语音处理
摘要
本文提出了 Cross Language Agent -- Simultaneous Interpretation (CLASI),一个高质量且贴近人类的同时语音翻译系统。灵感来自专业人类解释员,我们采用 novel data-driven read-write 策略来平衡翻译质量与延迟。为解决翻译领域术语的挑战,CLASI 采用多模态检索模块获取相关信息以增强翻译。基于 LLM,我们的方法可生成容错翻译,考虑输入音频、历史语境及检索信息。实验结果表明,我们的系统以显著优势超越其他系统。与专业人类解释员一致,我们采用更好的人类评估指标 valid information proportion (VIP),衡量信息成功传递给听众的程度。在真实场景中,由于言语常不连贯、非正式且模糊不清,CLASI 对中文到英文与英文到中文翻译方向分别实现 81.3% 与 78.0% 的 VIP。相比,领先的商业或开源系统仅实现 35.4% 与 41.6%。在极端困难数据集上,其他系统实现不到 13% 的 VIP,而 CLASI 仍可达 70%。
引用
@article{arxiv.2407.21646,
title = {Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent},
author = {Shanbo Cheng and Zhichao Huang and Tom Ko and Hang Li and Ningxin Peng and Lu Xu and Qini Zhang},
journal= {arXiv preprint arXiv:2407.21646},
year = {2024}
}
备注
Authors are listed in alphabetical order by last name. Demonstrations and human-annotated test sets are available at https://byteresearchcla.github.io/clasi