English

Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens

Sound 2026-01-21 v1 Audio and Speech Processing

Abstract

Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models are prone to hallucination errors, such as phoneme omissions and speaker inconsistency, which conventional error filtering based on non-intrusive speech quality metrics often fails to detect. To address this issue, we propose a non-intrusive method for filtering hallucination errors from discrete token-based GSE models. Our method leverages the log-probabilities of generated tokens as confidence scores to detect potential errors. Experimental results show that the confidence scores strongly correlate with a suite of intrusive SE metrics, and that our method effectively identifies hallucination errors missed by conventional filtering methods. Furthermore, we demonstrate the practical utility of our method: curating an in-the-wild TTS dataset with our confidence-based filtering improves the performance of subsequently trained TTS models.

Keywords

Cite

@article{arxiv.2601.12254,
  title  = {Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens},
  author = {Kazuki Yamauchi and Masato Murata and Shogo Seki},
  journal= {arXiv preprint arXiv:2601.12254},
  year   = {2026}
}

Comments

Accepted for ICASSP 2026

R2 v1 2026-07-01T09:09:14.828Z