中文

EnCLAP:结合神经音频编解码器与音频-文本联合嵌入的自动音频描述生成

音频与语音处理 2024-02-01 v1 人工智能 声音

摘要

我们提出了 EnCLAP,一种用于自动音频描述生成的新型框架。EnCLAP 采用了两个声学表示模型 EnCodec 和 CLAP,以及一个预训练语言模型 BART。我们还引入了一种称为掩码编解码建模的新训练目标,以提高预训练语言模型的声学感知能力。在 AudioCaps 和 Clotho 上的实验结果表明,我们的模型超越了基线模型的性能。源代码将发布于 https://github.com/jaeyeonkim99/EnCLAP 。在线演示可访问 https://huggingface.co/spaces/enclap-team/enclap 。

关键词

引用

@article{arxiv.2401.17690,
  title  = {EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning},
  author = {Jaeyeon Kim and Jaeyoon Jung and Jinjoo Lee and Sang Hoon Woo},
  journal= {arXiv preprint arXiv:2401.17690},
  year   = {2024}
}

备注

Accepted to ICASSP 2024