English
Related papers

Related papers: BLSP-Emo: Towards Empathetic Large Speech-Language…

200 papers

End-to-end speech translation models have become a new trend in research due to their potential of reducing error propagation. However, these models still suffer from the challenge of data scarcity. How to effectively use unlabeled or other…

Computation and Language · Computer Science 2021-06-21 Rong Ye , Mingxuan Wang , Lei Li

Understanding emotions and responding accordingly is one of the biggest challenges of dialog systems. This paper presents EmpTransfo, a multi-head Transformer architecture for creating an empathetic dialog system. EmpTransfo utilizes…

Computation and Language · Computer Science 2020-03-09 Rohola Zandie , Mohammad H. Mahoor

Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal emotion recognition capabilities, integrating multimodal cues from visual, acoustic, and linguistic contexts in the video to recognize human emotional states.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Liyun Zhang

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vision, speech, and…

Emotion recognition in conversation (ERC) aims to detect the emotion for each utterance in a given conversation. The newly proposed ERC models have leveraged pre-trained language models (PLMs) with the paradigm of pre-training and…

Computation and Language · Computer Science 2022-07-28 Jingjie Yi , Deqing Yang , Siyu Yuan , Caiyan Cao , Zhiyao Zhang , Yanghua Xiao

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

This paper builds upon an existing speech emotion recognition model by adding an additional LSTM layer to improve the accuracy and processing efficiency of emotion recognition from audio data. By capturing the long-term dependencies within…

Artificial Intelligence · Computer Science 2024-12-02 Xiaoran Yang , Shuhan Yu , Wenxi Xu

Language model pre-training has shown promising results in various downstream tasks. In this context, we introduce a cross-modal pre-trained language model, called Speech-Text BERT (ST-BERT), to tackle end-to-end spoken language…

Computation and Language · Computer Science 2021-04-13 Minjeong Kim , Gyuwan Kim , Sang-Woo Lee , Jung-Woo Ha

Previous works on emotion recognition in conversation (ERC) follow a two-step paradigm, which can be summarized as first producing context-independent features via fine-tuning pretrained language models (PLMs) and then analyzing contextual…

Computation and Language · Computer Science 2023-01-18 Xiangyu Qin , Zhiyu Wu , Jinshi Cui , Tingting Zhang , Yanran Li , Jian Luan , Bin Wang , Li Wang

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Xiaoxue Gao , Huayun Zhang , Nancy F. Chen

We propose a workflow for speech emotion recognition (SER) that combines pre-trained representations with automated hyperparameter optimisation (HPO). Using SpeechBrain wav2vec2-base model fine-tuned on IEMOCAP as the encoder, we compare…

Machine Learning · Computer Science 2025-10-09 Aryan Golbaghi , Shuo Zhou

We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationalised by the TRUST multi-agent large language model (LLM) pipeline. Using a Bundestag…

Artificial Intelligence · Computer Science 2026-05-22 Juergen Dietrich

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Haobin Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST-Tacotron framework. The style tokens are defined to present…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-27 Peng-fei Wu , Zhen-hua Ling , Li-juan Liu , Yuan Jiang , Hong-chuan Wu , Li-rong Dai

In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an…

Sound · Computer Science 2025-08-29 ChenYi Chua , JunKai Wong , Chengxin Chen , Xiaoxiao Miao

Multimodal Emotion Recognition (MER) aims to automatically identify and understand human emotional states by integrating information from various modalities. However, the scarcity of annotated multimodal data significantly hinders the…

Human-Computer Interaction · Computer Science 2024-09-11 Zhixian Zhao , Haifeng Chen , Xi Li , Dongmei Jiang , Lei Xie

Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is challenging to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-17 Ye Bai , Jiangyan Yi , Jianhua Tao , Zhengqi Wen , Zhengkun Tian , Shuai Zhang

Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach to tackle these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Aref Farhadipour , Hossein Ranjbar , Masoumeh Chapariniya , Teodora Vukovic , Sarah Ebling , Volker Dellwo

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to…

Human-Computer Interaction · Computer Science 2025-05-08 Zheng Lian , Haiyang Sun , Licai Sun , Haoyu Chen , Lan Chen , Hao Gu , Zhuofan Wen , Shun Chen , Siyuan Zhang , Hailiang Yao , Bin Liu , Rui Liu , Shan Liang , Ya Li , Jiangyan Yi , Jianhua Tao

We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history. Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and…