Auden-Voice: 通用语音编码器用于语音和语言理解
音频与语音处理
2025-11-20 v1 声音
摘要
人类语音编码了身份信息和伴随语言线索,然而大型音频语言模型 (LALM) 中的编码器很少在这两方面之间取得平衡。本文致力于构建通用语音编码器以捕捉细微的语音线索。通过全面评估,我们发现多任务训练可获得最均衡的表示,而对比语言音频预训练 (CLAP) 主要提高检索效果,却未提升伴随语言理解。我们的最终编码器 Auden-Voice 也在集成到大语言模型 (LLM) 时表现出色。代码和训练配方将随音频理解工具包 Auden 一起发布。
引用
@article{arxiv.2511.15145,
title = {Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding},
author = {Mingyue Huo and Wei-Cheng Tseng and Yiwen Shao and Hao Zhang and Dong Yu},
journal= {arXiv preprint arXiv:2511.15145},
year = {2025}
}
备注
Submitted to ICASSP2026