English

Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora

Audio and Speech Processing 2025-07-03 v1 Sound

Abstract

Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples tailored to target audiences would enable users to adjust their speaking style and voice quality, facilitating smoother communication. To this end, we propose a voice conversion method that controls the likability of input speech while preserving both speaker identity and linguistic content. To improve training data scalability, we train a likability predictor on an existing voice likability dataset and employ it to automatically annotate a large speech synthesis corpus with likability ratings. Experimental evaluations reveal a significant correlation between the predictor's outputs and human-provided likability ratings. Subjective and objective evaluations further demonstrate that the proposed approach effectively controls voice likability while preserving both speaker identity and linguistic content.

Keywords

Cite

@article{arxiv.2507.01356,
  title  = {Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora},
  author = {Hitoshi Suda and Shinnosuke Takamichi and Satoru Fukayama},
  journal= {arXiv preprint arXiv:2507.01356},
  year   = {2025}
}

Comments

Accepted at Interspeech 2025