基于隐式与显式音素条件的语音音素与韵律注释
计算机视觉与模式识别
2025-07-16 v3
摘要
我们提出一种模型,以与音素一致的方式获取语音的音素和韵律标签。不同于以往仅通过对预训练ASR模型进行微调以获得标签的方法,本文提出的模型通过两种方法将标签生成条件化于相应音素:1)通过使用预训练BERT特征的提示编码器引入隐式音素条件;2)在推理时对与音素不一致的标签假设进行显式修剪。这些方法使得能够获取语音、标签和音素的平行数据,这种数据适用于各种下游任务,如文本转语音和从文本进行方言估计。实验表明,所提方法显著提高了音素与预测标签之间的一致性。进一步的实验在方言估计任务上确认,所提方法创建的平行数据有效提升了估计准确度。
引用
@article{arxiv.2506.04526,
title = {EECD-Net: Energy-Efficient Crack Detection with Spiking Neural Networks and Gated Attention},
author = {Shuo Zhang},
journal= {arXiv preprint arXiv:2506.04526},
year = {2025}
}
备注
After further careful review and additional checks, we have identified multiple issues in our experimental results and data analysis that significantly affect the validity and reliability of our findings. We believe that these issues are substantial enough to compromise the scientific integrity of the manuscript