English

Improving speaker de-identification with functional data analysis of f0 trajectories

Sound 2022-04-01 v1 Computation and Language Audio and Speech Processing

Abstract

Due to a constantly increasing amount of speech data that is stored in different types of databases, voice privacy has become a major concern. To respond to such concern, speech researchers have developed various methods for speaker de-identification. The state-of-the-art solutions utilize deep learning solutions which can be effective but might be unavailable or impractical to apply for, for example, under-resourced languages. Formant modification is a simpler, yet effective method for speaker de-identification which requires no training data. Still, remaining intonational patterns in formant-anonymized speech may contain speaker-dependent cues. This study introduces a novel speaker de-identification method, which, in addition to simple formant shifts, manipulates f0 trajectories based on functional data analysis. The proposed speaker de-identification method will conceal plausibly identifying pitch characteristics in a phonetically controllable manner and improve formant-based speaker de-identification up to 25%.

Keywords

Cite

@article{arxiv.2203.16738,
  title  = {Improving speaker de-identification with functional data analysis of f0 trajectories},
  author = {Lauri Tavi and Tomi Kinnunen and Rosa González Hautamäki},
  journal= {arXiv preprint arXiv:2203.16738},
  year   = {2022}
}

Comments

Accepted to Speech Communication. March 2022

R2 v1 2026-06-24T10:32:45.966Z