English

Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis

Computer Vision and Pattern Recognition 2025-11-10 v1

Abstract

We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle distribution shifts between clean and TTS-predicted features, we adopt a two-stage training: pretraining on Wav2Vec2 embeddings and finetuning on TTS outputs. This enables tight audio-visual alignment, preserves speaker identity, and produces natural, expressive speech and synchronized facial motion without ground-truth audio at inference. Experiments show that conditioning on TTS-predicted latent features outperforms cascaded pipelines, improving both lip-sync and visual realism.

Keywords

Cite

@article{arxiv.2511.05432,
  title  = {Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis},
  author = {Dogucan Yaman and Seymanur Akti and Fevziye Irem Eyiokur and Alexander Waibel},
  journal= {arXiv preprint arXiv:2511.05432},
  year   = {2025}
}
R2 v1 2026-07-01T07:26:31.495Z