English

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

Computation and Language 2025-08-29 v2 Computer Vision and Pattern Recognition Multimedia Sound Audio and Speech Processing

Abstract

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems.

Keywords

Cite

@article{arxiv.2508.16188,
  title  = {Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation},
  author = {Weiting Tan and Jiachen Lian and Hirofumi Inaguma and Paden Tomasello and Philipp Koehn and Xutai Ma},
  journal= {arXiv preprint arXiv:2508.16188},
  year   = {2025}
}

Comments

EMNLP 2025 (Findings)

R2 v1 2026-07-01T05:01:22.673Z