English

On The Open Prompt Challenge In Conditional Audio Generation

Sound 2023-11-03 v1 Computation and Language Audio and Speech Processing

Abstract

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified when compared to text descriptions used to train TTA models. In this work, we treat TTA models as a ``blackbox'' and address the user prompt challenge with two key insights: (1) User prompts are generally under-specified, leading to a large alignment gap between user prompts and training prompts. (2) There is a distribution of audio descriptions for which TTA models are better at generating higher quality audio, which we refer to as ``audionese''. To this end, we rewrite prompts with instruction-tuned models and propose utilizing text-audio alignment as feedback signals via margin ranking learning for audio improvements. On both objective and subjective human evaluations, we observed marked improvements in both text-audio alignment and music audio quality.

Keywords

Cite

@article{arxiv.2311.00897,
  title  = {On The Open Prompt Challenge In Conditional Audio Generation},
  author = {Ernie Chang and Sidd Srinivasan and Mahi Luthra and Pin-Jie Lin and Varun Nagaraja and Forrest Iandola and Zechun Liu and Zhaoheng Ni and Changsheng Zhao and Yangyang Shi and Vikas Chandra},
  journal= {arXiv preprint arXiv:2311.00897},
  year   = {2023}
}

Comments

5 pages, 3 figures, 4 tables

R2 v1 2026-06-28T13:09:09.358Z