EnCLAP:结合神经音频编解码器与音频-文本联合嵌入的自动音频描述生成
音频与语音处理
2024-02-01 v1 人工智能
声音
摘要
我们提出了 EnCLAP,一种用于自动音频描述生成的新型框架。EnCLAP 采用了两个声学表示模型 EnCodec 和 CLAP,以及一个预训练语言模型 BART。我们还引入了一种称为掩码编解码建模的新训练目标,以提高预训练语言模型的声学感知能力。在 AudioCaps 和 Clotho 上的实验结果表明,我们的模型超越了基线模型的性能。源代码将发布于 https://github.com/jaeyeonkim99/EnCLAP 。在线演示可访问 https://huggingface.co/spaces/enclap-team/enclap 。
引用
@article{arxiv.2401.17690,
title = {EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning},
author = {Jaeyeon Kim and Jaeyoon Jung and Jinjoo Lee and Sang Hoon Woo},
journal= {arXiv preprint arXiv:2401.17690},
year = {2024}
}
备注
Accepted to ICASSP 2024