English

Reference-aware SFM layers for intrusive intelligibility prediction

Audio and Speech Processing 2025-09-23 v1 Sound

Abstract

Intrusive speech-intelligibility predictors that exploit explicit reference signals are now widespread, yet they have not consistently surpassed non-intrusive systems. We argue that a primary cause is the limited exploitation of speech foundation models (SFMs). This work revisits intrusive prediction by combining reference conditioning with multi-layer SFM representations. Our final system achieves RMSE 22.36 on the development set and 24.98 on the evaluation set, ranking 1st on CPC3. These findings provide practical guidance for constructing SFM-based intrusive intelligibility predictors.

Keywords

Cite

@article{arxiv.2509.17270,
  title  = {Reference-aware SFM layers for intrusive intelligibility prediction},
  author = {Hanlin Yu and Haoshuai Zhou and Boxuan Cao and Changgeng Mo and Linkai Li and Shan X. Wang},
  journal= {arXiv preprint arXiv:2509.17270},
  year   = {2025}
}

Comments

Preprint; submitted to ICASSP 2026. 5 pages. CPC3 system: Dev RMSE 22.36, Eval RMSE 24.98 (ranked 1st)

R2 v1 2026-07-01T05:48:39.780Z