用于自动发声模式分类的数据集
摘要
Complete Vocal Technique (CVT) 是过去几十年由 Cathrin Sadolin 等人发展起来的一门歌唱学派。CVT 将声音的使用归类为所谓的发声模式,即 Neutral、Curbing、Overdrive 和 Edge。了解所期望的发声模式对歌唱学生很有帮助。因此,发声模式的自动分类对于技术辅助的歌唱教学可能具有重要意义。此前,自动分类发声模式的尝试未取得重大成功,可能是由于缺乏数据。因此,我们录制了一个新颖的发声模式数据集,包含来自四名歌手的持续元音,其中三名为具有五年以上 CVT 经验的专业歌手。该数据集涵盖了受试者的整个音域,共计 3,752 个独特样本。通过使用四个麦克风从而提供自然的数据增强,该数据集合并后包含超过 13,000 个样本。标注由三名具有 CVT 经验的标注员完成,每人提供一份独立标注。合并的标注以及三份独立标注随发布的数据集一起提供。此外,我们提供了一些基线分类结果。在 5 折交叉验证中,使用 ResNet18 取得了 81.3% 的最佳平衡准确率。该数据集可在 https://zenodo.org/records/14276415 下载。
引用
@article{arxiv.2601.18339,
title = {A Dataset for Automatic Vocal Mode Classification},
author = {Reemt Hinrichs and Sonja Stephan and Alexander Lange and Jörn Ostermann},
journal= {arXiv preprint arXiv:2601.18339},
year = {2026}
}
备注
Extended manuscript of our Article in the proceedings of the EvoMUSART 2026: 15th International Conference on Artificial Intelligence in Music, Sound, Art and Design; Tiny corrigendum to v1, where the pitch distribution showed an incorrect F1. The truely lowest note of the dataset is a B1