English
Related papers

Related papers: MFSN: Multi-perspective Fusion Search Network For …

200 papers

Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training. Recently, segment-based approaches for SER have been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-31 Shuiyang Mao , P. C. Ching , Tan Lee

Multi-view sequential learning is a fundamental problem in machine learning dealing with multi-view sequences. In a multi-view sequence, there exists two forms of interactions between different views: view-specific interactions and…

Machine Learning · Computer Science 2018-02-06 Amir Zadeh , Paul Pu Liang , Navonil Mazumder , Soujanya Poria , Erik Cambria , Louis-Philippe Morency

Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech…

Sound · Computer Science 2025-07-11 Zhao Ren , Rathi Adarshi Rammohan , Kevin Scheck , Sheng Li , Tanja Schultz

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP)…

Sound · Computer Science 2025-07-08 Jiacheng Shi , Yanfu Zhang , Ye Gao

A graph neural network (GNN) for image understanding based on multiple cues is proposed in this paper. Compared to traditional feature and decision fusion approaches that neglect the fact that features can interact and exchange information,…

Computer Vision and Pattern Recognition · Computer Science 2020-03-02 Xin Guo , Luisa F. Polania , Bin Zhu , Charles Boncelet , Kenneth E. Barner

The purpose of emotion recognition in conversation (ERC) is to identify the emotion category of an utterance based on contextual information. Previous ERC methods relied on simple connections for cross-modal fusion and ignored the…

Computation and Language · Computer Science 2024-05-29 Haoxiang Shi , Xulong Zhang , Ning Cheng , Yong Zhang , Jun Yu , Jing Xiao , Jianzong Wang

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Fueled by recent advances of self-supervised models, pre-trained speech representations proved effective for the downstream speech emotion recognition (SER) task. Most prior works mainly focus on exploiting pre-trained representations and…

Sound · Computer Science 2023-03-02 Siyuan Shen , Feng Liu , Aimin Zhou

Humans are able to comprehend information from multiple domains for e.g. speech, text and visual. With advancement of deep learning technology there has been significant improvement of speech recognition. Recognizing emotion from speech is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-16 Mandeep Singh , Yuan Fang

While speech emotion recognition (SER) research has made significant progress, achieving generalization across various corpora continues to pose a problem. We propose a novel domain adaptation technique that embodies a multitask framework…

Computation and Language · Computer Science 2023-10-10 Chung-Soo Ahn , Jagath C. Rajapakse , Rajib Rana

Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results…

Computation and Language · Computer Science 2024-05-31 Zhuoyuan Mao , Chenhui Chu , Sadao Kurohashi

Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models. The majority of the studies focus on settings where both modalities are available in training and…

Computation and Language · Computer Science 2019-06-26 Gustavo Aguilar , Viktor Rozgić , Weiran Wang , Chao Wang

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Esther Sun , Abinay Reddy Naini , Carlos Busso

Speech emotion recognition (SER) is an important part of human-computer interaction, receiving extensive attention from both industry and academia. However, the current research field of SER has long suffered from the following problems: 1)…

Sound · Computer Science 2024-06-12 Ziyang Ma , Mingjie Chen , Hezhao Zhang , Zhisheng Zheng , Wenxi Chen , Xiquan Li , Jiaxin Ye , Xie Chen , Thomas Hain

Generative adversarial networks (GANs) have shown potential in learning emotional attributes and generating new data samples. However, their performance is usually hindered by the unavailability of larger speech emotion recognition (SER)…

Sound · Computer Science 2020-07-28 Siddique Latif , Muhammad Asim , Rajib Rana , Sara Khalifa , Raja Jurdak , Björn W. Schuller

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible…

Computation and Language · Computer Science 2026-03-18 Yuan Ge , Haishu Zhao , Aokai Hao , Junxiang Zhang , Bei Li , Xiaoqian Liu , Chenglong Wang , Jianjin Wang , Bingsen Zhou , Bingyu Liu , Jingbo Zhu , Zhengtao Yu , Tong Xiao

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges when collecting sensitive emotional states. We introduce…

Sentiment analysis is known as one of the most crucial tasks in the field of natural language processing and Convolutional Neural Network (CNN) is one of those prominent models that is commonly used for this aim. Although convolutional…

Computation and Language · Computer Science 2021-02-24 Hossein Sadr , Mozhdeh Nazari Solimandarabi , Mir Mohsen Pedram , Mohammad Teshnehlab

In this paper, we propose a novel deep inductive transfer learning framework, named feature distribution adaptation network, to tackle the challenging multi-modal speech emotion recognition problem. Our method aims to use deep transfer…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Shaokai Li , Yixuan Ji , Peng Song , Haoqin Sun , Wenming Zheng

Multimodal Emotion Recognition (MER) aims to automatically identify and understand human emotional states by integrating information from various modalities. However, the scarcity of annotated multimodal data significantly hinders the…

Human-Computer Interaction · Computer Science 2024-09-11 Zhixian Zhao , Haifeng Chen , Xi Li , Dongmei Jiang , Lei Xie