English
Related papers

Related papers: Emotion Recognition in Speech using Cross-Modal Tr…

200 papers

To train machine learning algorithms to predict emotional expressions in terms of arousal and valence, annotated datasets are needed. However, as different people perceive others' emotional expressions differently, their annotations are…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-14 Navin Raj Prabhu , Nale Lehmann-Willenbrock , Timo Gerkman

Emotion recognition models using audio input data can enable the development of interactive systems with applications in mental healthcare, marketing, gaming, and social media analysis. While the field of affective computing using audio…

Sound · Computer Science 2023-07-25 Peranut Nimitsurachat , Peter Washington

Key features of mental illnesses are reflected in speech. Our research focuses on designing a multimodal deep learning structure that automatically extracts salient features from recorded speech samples for predicting various mental…

Machine Learning · Computer Science 2020-04-15 Habibeh Naderi , Behrouz Haji Soleimani , Stan Matwin

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study…

Sound · Computer Science 2025-08-22 Ilya Fedorov , Dmitry Korobchenko

While the performance of cross-lingual TTS based on monolingual corpora has been significantly improved recently, generating cross-lingual speech still suffers from the foreign accent problem, leading to limited naturalness. Besides,…

Sound · Computer Science 2023-09-06 Tao Li , Chenxu Hu , Jian Cong , Xinfa Zhu , Jingbei Li , Qiao Tian , Yuping Wang , Lei Xie

Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-15 Huang-Cheng Chou , Haibin Wu , Hung-yi Lee , Chi-Chun Lee

We introduce a framework that recommends music based on the emotions of speech. In content creation and daily life, speech contains information about human emotions, which can be enhanced by music. Our framework focuses on a cross-domain…

Sound · Computer Science 2023-03-21 SeungHeon Doh , Minz Won , Keunwoo Choi , Juhan Nam

Emotion Recognition in Conversation is a core component of affective computing, while current resources of sign language emotion datasets primarily focus on isolated sentences and lack conversational context. Models trained exclusively on…

Computation and Language · Computer Science 2026-05-25 Yusong Wang , Keyu Mao , Takao Obi , Minghao Shao , Kotaro Funakoshi

Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold standard references. In addition, there is much variability from input speech signals, human subjective perception of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-13 Haoqi Li , Ming Tu , Jing Huang , Shrikanth Narayanan , Panayiotis Georgiou

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network…

Sound · Computer Science 2021-02-12 Kun Zhou , Berrak Sisman , Rui Liu , Haizhou Li

Learning speaker turn embeddings has shown considerable improvement in situations where conventional speaker modeling approaches fail. However, this improvement is relatively limited when compared to the gain observed in face embedding…

Computer Vision and Pattern Recognition · Computer Science 2017-07-11 Nam Le , Jean-Marc Odobez

For speech emotion datasets, it has been difficult to acquire large quantities of reliable data and acted emotions may be over the top compared to less expressive emotions displayed in everyday life. Lately, larger datasets with natural…

Computation and Language · Computer Science 2022-07-06 Rosanna Milner , Md Asif Jalal , Raymond W. M. Ng , Thomas Hain

When recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-04 Sofoklis Kakouros , Themos Stafylakis , Ladislav Mosner , Lukas Burget

Embeddings play an important role in end-to-end solutions for multi-modal language processing problems. Although there has been some effort to understand the properties of single-modality embedding spaces, particularly that of text, their…

Computation and Language · Computer Science 2023-01-20 Muhammad Huzaifah , Ivan Kukanov

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of…

Sound · Computer Science 2023-10-09 Yuke Li , Xinfa Zhu , Yi Lei , Hai Li , Junhui Liu , Danming Xie , Lei Xie

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is…

Sound · Computer Science 2022-06-28 Yihan Wu , Xi Wang , Shaofei Zhang , Lei He , Ruihua Song , Jian-Yun Nie

Speech emotion recognition plays an important role in building more intelligent and human-like agents. Due to the difficulty of collecting speech emotional data, an increasingly popular solution is leveraging a related and rich source…

Machine Learning · Computer Science 2019-02-15 Hao Zhou , Ke Chen

In this paper, we describe our algorithmic approach, which was used for submissions in the fifth Emotion Recognition in the Wild (EmotiW 2017) group-level emotion recognition sub-challenge. We extracted feature vectors of detected faces…

Computer Vision and Pattern Recognition · Computer Science 2017-11-07 Alexandr G. Rassadin , Alexey S. Gruzdev , Andrey V. Savchenko