English

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Sound 2025-08-12 v1 Artificial Intelligence Computation and Language Audio and Speech Processing

Abstract

Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.

Keywords

Cite

@article{arxiv.2508.06890,
  title  = {Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody},
  author = {Jinsung Yoon and Wooyeol Jeong and Jio Gim and Young-Joo Suh},
  journal= {arXiv preprint arXiv:2508.06890},
  year   = {2025}
}

Comments

Accepted at ASRU 2025

R2 v1 2026-07-01T04:42:21.312Z