English

Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study

Sound 2024-07-30 v3 Artificial Intelligence Multimedia Audio and Speech Processing

Abstract

Deep neural networks have recently achieved breakthroughs in sound generation. Despite the outstanding sample quality, current sound generation models face issues on small-scale datasets (e.g., overfitting), significantly limiting performance. In this paper, we make the first attempt to investigate the benefits of pre-training on sound generation with AudioLDM, the cutting-edge model for audio generation, as the backbone. Our study demonstrates the advantages of the pre-trained AudioLDM, especially in data-scarcity scenarios. In addition, the baselines and evaluation protocol for sound generation systems are not consistent enough to compare different studies directly. Aiming to facilitate further study on sound generation tasks, we benchmark the sound generation task on various frequently-used datasets. We hope our results on transfer learning and benchmarks can provide references for further research on conditional sound generation.

Keywords

Cite

@article{arxiv.2303.03857,
  title  = {Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study},
  author = {Yi Yuan and Haohe Liu and Jinhua Liang and Xubo Liu and Mark D. Plumbley and Wenwu Wang},
  journal= {arXiv preprint arXiv:2303.03857},
  year   = {2024}
}

Comments

Updated for EUSIPCO 2023 proceedings version

R2 v1 2026-06-28T09:05:25.464Z