PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
Abstract
Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack rich variations and diversity, resulting in a drop in TTA model performance when faced with complex text. To address this issue, we propose a method called Portable Plug-in Prompt Refiner, which utilizes rich knowledge about textual descriptions inherent in large language models to effectively enhance the robustness of TTA acoustic models without altering the acoustic training set. Furthermore, a Chain-of-Thought that mimics human verification is introduced to enhance the accuracy of audio descriptions, thereby improving the accuracy of generated content in practical applications. The experiments show that our method achieves a state-of-the-art Inception Score (IS) of 8.72, surpassing AudioGen, AudioLDM and Tango.
Cite
@article{arxiv.2406.04683,
title = {PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation},
author = {Shuchen Shi and Ruibo Fu and Zhengqi Wen and Jianhua Tao and Tao Wang and Chunyu Qiang and Yi Lu and Xin Qi and Xuefei Liu and Yukun Liu and Yongwei Li and Zhiyong Wang and Xiaopeng Wang},
journal= {arXiv preprint arXiv:2406.04683},
year = {2024}
}
Comments
accepted by INTERSPEECH2024