自动痴呆评估中的常见陷阱与限制
音频与语音处理
2025-08-07 v1
摘要
当前关于基于语音的痴呆评估的工作要么 focuses on 提取特征来预测评估量表,要么 focuses on 自动化现有测试程序的自动化。大多数研究无条件地使用公共数据,很少进行详细的错误分析,主要关注数值性能。我们对一种自动化标准化痴呆评估——Syndrom-Kurz-Test 进行了深入分析。我们发现尽管总体与人类标注员之间存在较高相关性,但由于某些 artifact,观察到对严重受损个体的相关性较高,而对健康或轻度受损个体则不然。随着认知功能下降,语音产出减少,导致当 test scoring relies on 单词命名时产生乐观的相关性。 Depending on test design, fallback handling 引入了进一步的偏见,偏向特定群体。这些陷阱独立于数据集中的群体分布,require 对目标群体进行差异化分析。
关键词
引用
@article{arxiv.2508.04512,
title = {Pitfalls and Limits in Automatic Dementia Assessment},
author = {Franziska Braun and Christopher Witzl and Andreas Erzigkeit and Hartmut Lehfeld and Thomas Hillemacher and Tobias Bocklet and Korbinian Riedhammer},
journal= {arXiv preprint arXiv:2508.04512},
year = {2025}
}
备注
Accepted at INTERSPEECH 2025