English

EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

Computation and Language 2024-10-15 v2 Sound Audio and Speech Processing

Abstract

We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.

Keywords

Cite

@article{arxiv.2312.14069,
  title  = {EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models},
  author = {Maureen de Seyssel and Antony D'Avirro and Adina Williams and Emmanuel Dupoux},
  journal= {arXiv preprint arXiv:2312.14069},
  year   = {2024}
}

Comments

Accepted at EMNLP 2024 (Main)