English
Related papers

Related papers: Optimizing Speech Emotion Recognition using Manta-…

200 papers

We propose a workflow for speech emotion recognition (SER) that combines pre-trained representations with automated hyperparameter optimisation (HPO). Using SpeechBrain wav2vec2-base model fine-tuned on IEMOCAP as the encoder, we compare…

Machine Learning · Computer Science 2025-10-09 Aryan Golbaghi , Shuo Zhou

Sentiment Analysis refers to the study of systematically extracting the meaning of subjective text . When analysing sentiments from the subjective text using Machine Learning techniques,feature extraction becomes a significant part. We…

Computation and Language · Computer Science 2019-06-05 Avinash Madasu , Sivasankar E

Speech Emotion Recognition (SER) is to recognize human emotions in a natural verbal interaction scenario with machines, which is considered as a challenging problem due to the ambiguous human emotions. Despite the recent progress in SER,…

Computation and Language · Computer Science 2023-05-11 Lei Kang , Lichao Zhang , Dazhi Jiang

Categorical speech emotion recognition is typically performed as a sequence-to-label problem, i.e., to determine the discrete emotion label of the input utterance as a whole. One of the main challenges in practice is that most of the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Shuiyang Mao , P. C. Ching , C. -C. Jay Kuo , Tan Lee

Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed…

Computation and Language · Computer Science 2026-03-24 Florian Lecourt , Madalina Croitoru , Konstantin Todorov

We examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptors and their…

This paper proposes a Residual Convolutional Neural Network (ResNet) based on speech features and trained under Focal Loss to recognize emotion in speech. Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCCs)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-16 Suraj Tripathi , Abhay Kumar , Abhiram Ramesh , Chirag Singh , Promod Yenigalla

This paper focuses on finding suitable features to robustly recognize emotions and evaluate customer satisfaction from speech in real acoustic scenarios. The classification of emotions is based on standard and well-known corpora and the…

Sound · Computer Science 2021-08-30 Luis Felipe Parra-Gallego , Juan Rafael Orozco-Arroyave

This paper presents a widespread analysis of affective vocal expression classification systems. In this study, state-of-the-art acoustic features are compared to two novel affective vocal prints for the detection of emotional states: the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-07 V. Vieira , R. Coelho , F. Assis

We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more…

Computation and Language · Computer Science 2022-01-31 Ayoub Ghriss , Bo Yang , Viktor Rozgic , Elizabeth Shriberg , Chao Wang

Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial intelligence (AI). Our proposed DCRF-BiLSTM model is used to recognize seven emotions: neutral,…

Sound · Computer Science 2026-01-15 Shahana Yasmin Chowdhury , Bithi Banik , Md Tamjidul Hoque , Shreya Banerjee

In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an…

Sound · Computer Science 2025-08-29 ChenYi Chua , JunKai Wong , Chengxin Chen , Xiaoxiao Miao

In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech…

Machine Learning · Computer Science 2024-03-28 Leonardo Pepino , Pablo Riera , Luciana Ferrer , Agustin Gravano

In this work, we focus on the detection of depression through speech analysis. Previous research has widely explored features extracted from pre-trained models (PTMs) primarily trained for paralinguistic tasks. Although these features have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Orchid Chetia Phukan , Sarthak Jain , Shubham Singh , Muskaan Singh , Arun Balaji Buduru , Rajesh Sharma

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional…

Artificial Intelligence · Computer Science 2024-11-05 Zebang Cheng , Zhi-Qi Cheng , Jun-Yan He , Jingdong Sun , Kai Wang , Yuxiang Lin , Zheng Lian , Xiaojiang Peng , Alexander Hauptmann

Music emotion recognition (MER) is usually regarded as a multi-label tagging task, and each segment of music can inspire specific emotion tags. Most researchers extract acoustic features from music and explore the relations between these…

Multimedia · Computer Science 2017-04-20 Xin Liu , Qingcai Chen , Xiangping Wu , Yan Liu , Yang Liu

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Soumya Dutta , Sriram Ganapathy

Emotion recognition is involved in several real-world applications. With an increase in available modalities, automatic understanding of emotions is being performed more accurately. The success in Multimodal Emotion Recognition (MER),…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Riccardo Franceschini , Enrico Fini , Cigdem Beyan , Alessandro Conti , Federica Arrigoni , Elisa Ricci

An algorithm involving Mel-Frequency Cepstral Coefficients (MFCCs) is provided to perform signal feature extraction for the task of speaker accent recognition. Then different classifiers are compared based on the MFCC feature. For each…

Sound · Computer Science 2015-02-02 Zichen Ma , Ernest Fokoue

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman