English

Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning

Audio and Speech Processing 2020-09-25 v1 Computation and Language Machine Learning Sound

Abstract

The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, dataaugmentation, multi-task learning, and post-processing, for audiocaptioning. The system received the highest evaluation scores, butwhich of the individual elements most fully contributed to its perfor-mance has not yet been clarified. Here, to asses their contributions,we first conducted an element-wise ablation study on our systemto estimate to what extent each element is effective. We then con-ducted a detailed module-wise ablation study to further clarify thekey processing modules for improving accuracy. The results showthat data augmentation and post-processing significantly improvethe score in our system. In particular, mix-up data augmentationand beam search in post-processing improve SPIDEr by 0.8 and 1.6points, respectively.

Keywords

Cite

@article{arxiv.2009.11436,
  title  = {Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning},
  author = {Daiki Takeuchi and Yuma Koizumi and Yasunori Ohishi and Noboru Harada and Kunio Kashino},
  journal= {arXiv preprint arXiv:2009.11436},
  year   = {2020}
}

Comments

Accepted to DCASE2020 Workshop

R2 v1 2026-06-23T18:45:25.665Z