English

Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition

Artificial Intelligence 2026-02-05 v2

Abstract

Large audio-language models (LALMs) exhibit strong zero-shot performance across speech tasks but struggle with speech emotion recognition (SER) due to weak paralinguistic modeling and limited cross-modal reasoning. We propose Compositional Chain-of-Thought Prompting for Emotion Reasoning (CCoT-Emo), a framework that introduces structured Emotion Graphs (EGs) to guide LALMs in emotion inference without fine-tuning. Each EG encodes seven acoustic features (e.g., pitch, speech rate, jitter, shimmer), textual sentiment, keywords, and cross-modal associations. Embedded into prompts, EGs provide interpretable and compositional representations that enhance LALM reasoning. Experiments across SER benchmarks show that CCoT-Emo outperforms prior SOTA and improves accuracy over zero-shot baselines.

Keywords

Cite

@article{arxiv.2509.25458,
  title  = {Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition},
  author = {Jiacheng Shi and Hongfei Du and Y. Alicia Hong and Ye Gao},
  journal= {arXiv preprint arXiv:2509.25458},
  year   = {2026}
}

Comments

Accepted to ICASSP 2026

R2 v1 2026-07-01T06:06:09.115Z