English

Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

Computation and Language 2025-08-12 v1 Artificial Intelligence Audio and Speech Processing

Abstract

Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.

Keywords

Cite

@article{arxiv.2508.07273,
  title  = {Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models},
  author = {Qiongqiong Wang and Hardik B. Sailor and Jeremy H. M. Wong and Tianchi Liu and Shuo Sun and Wenyu Zhang and Muhammad Huzaifah and Nancy Chen and Ai Ti Aw},
  journal= {arXiv preprint arXiv:2508.07273},
  year   = {2025}
}

Comments

Accepted at (ASRU 2025) 2025 IEEE Automatic Speech Recognition and Understanding Workshop

R2 v1 2026-07-01T04:42:59.593Z