English

Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning

Computation and Language 2026-03-18 v1 Artificial Intelligence

Abstract

Speech large language models (LLMs) observe paralinguistic cues such as prosody, emotion, and non-verbal sounds--crucial for intent understanding. However, leveraging these cues faces challenges: limited training data, annotation difficulty, and models exploiting lexical shortcuts over paralinguistic signals. We propose multi-task reinforcement learning (RL) with chain-of-thought prompting that elicits explicit affective reasoning. To address data scarcity, we introduce a paralinguistics-aware speech LLM (PALLM) that jointly optimizes sentiment classification from audio and paralinguistics-aware response generation via a two-stage pipeline. Experiments demonstrate that our approach improves paralinguistics understanding over both supervised baselines and strong proprietary models (Gemini-2.5-Pro, GPT-4o-audio) by 8-12% on Expresso, IEMOCAP, and RAVDESS. The results show that modeling paralinguistic reasoning with multi-task RL is crucial for building emotionally intelligent speech LLMs.

Keywords

Cite

@article{arxiv.2603.15981,
  title  = {Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning},
  author = {Jingxiang Chen and Minseok Kim and Seong-Gyun Leem and Yin Huang and Rashi Rungta and Zhicheng Ouyang and Haibin Wu and Surya Teja Appini and Ankur Bansal and Yang Bai and Yue Liu and Florian Metze and Ahmed A Aly and Anuj Kumar and Ariya Rastrow and Zhaojiang Lin},
  journal= {arXiv preprint arXiv:2603.15981},
  year   = {2026}
}
R2 v1 2026-07-01T11:23:19.684Z