English

An Audio Language Model-Based Voice Concept Bottleneck Framework for Interpretable Health Assessment

Audio and Speech Processing 2026-07-18 v1

Abstract

Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited. In this study, we propose a voice concept bottleneck framework for interpretable health assessment using an audio language model (ALM). The ALM is fine-tuned on a voice quality assessment dataset to enhance its understanding of voice concepts and serves as an independent concept extractor, producing discrete, interpretable scores for a lightweight downstream classifier. The discrete concept scores provide intuitive interpretation, while the lightweight classifier facilitates post-hoc interpretability analyses. Results on depression and dysarthria assessment tasks demonstrate that the proposed framework can flexibly adapt voice concepts to different health conditions and consistently outperforms openSMILE-based and self-supervised speech model-based baselines.

Cite

@article{arxiv.2607.16967,
  title  = {An Audio Language Model-Based Voice Concept Bottleneck Framework for Interpretable Health Assessment},
  author = {Yu-Wen Chen and Julia Hirschberg},
  journal= {arXiv preprint arXiv:2607.16967},
  year   = {2026}
}