中文

基于Wav2vec的构音障碍语音检测与严重程度分级

音频与语音处理 2023-10-18 v2 计算与语言 机器学习 声音 信号处理

摘要

直接从声学语音信号自动检测构音障碍并进行严重程度分级可用作医学诊断中的工具。本工作中,预训练的wav2vec 2.0模型被研究作为特征提取器,以构建构音障碍语音的检测与严重程度分级系统。实验使用广泛使用的UA-speech数据库进行。在检测实验中,结果显示使用wav2vec模型第一层的嵌入取得了最佳性能,相较最佳基线特征(谱图)准确率绝对提升1.23%。在所研究的严重程度分级任务中,结果显示最终层的嵌入相较最佳基线特征(梅尔频率倒谱系数)准确率绝对提升10.62%。

关键词

引用

@article{arxiv.2309.14107,
  title  = {Wav2vec-based Detection and Severity Level Classification of Dysarthria from Speech},
  author = {Farhad Javanmardi and Saska Tirronen and Manila Kodali and Sudarsana Reddy Kadiri and Paavo Alku},
  journal= {arXiv preprint arXiv:2309.14107},
  year   = {2023}
}

备注

copyright 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works