English
Related papers

Related papers: EmoTech: A Multi-modal Speech Emotion Recognition …

200 papers

Compared with facial emotion recognition on categorical model, the dimensional emotion recognition can describe numerous emotions of the real world more accurately. Most prior works of dimensional emotion estimation only considered…

Computer Vision and Pattern Recognition · Computer Science 2018-11-30 Xiaohua Wang , Muzi Peng , Lijuan Pan , Min Hu , Chunhua Jin , Fuji Ren

There are a variety of features of the human voice that can be classified as pitch, timbre, loudness, and vocal tone. It is observed in numerous incidents that human expresses their feelings using different vocal qualities when they are…

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

Computation and Language · Computer Science 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu

Decoding the speech signal that a person is listening to from the human brain via electroencephalography (EEG) can help us understand how our auditory system works. Linear models have been used to reconstruct the EEG from speech or vice…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-18 Mohammad Jalilpour Monesi , Bernd Accou , Tom Francart , Hugo Van Hamme

Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Dung Nguyen , Duc Thanh Nguyen , Rui Zeng , Thanh Thi Nguyen , Son N. Tran , Thin Nguyen , Sridha Sridharan , Clinton Fookes

Recent analysis on speech emotion recognition has made considerable advances with the use of MFCCs spectrogram features and the implementation of neural network approaches such as convolutional neural networks (CNNs). Capsule networks…

Sound · Computer Science 2021-12-28 Ismail Shahin , Noor Hindawi , Ali Bou Nassif , Adi Alhudhaif , Kemal Polat

Visual speech recognition models traditionally consist of two stages, feature extraction and classification. Several deep learning approaches have been recently presented aiming to replace the feature extraction stage by automatically…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Stavros Petridis , Yujiang Wang , Pingchuan Ma , Zuwei Li , Maja Pantic

Research in automatic affect recognition has seldom addressed the issue of computational resource utilization. With the advent of ambient intelligence technology which employs a variety of low-power, resource-constrained devices, this issue…

Machine Learning · Computer Science 2020-06-01 Fasih Haider , Senja Pollak , Pierre Albert , Saturnino Luz

General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language…

Computation and Language · Computer Science 2020-11-03 Aparna Khare , Srinivas Parthasarathy , Shiva Sundaram

Multi-modal conversation emotion recognition (MCER) aims to recognize and track the speaker's emotional state using text, speech, and visual information in the conversation scene. Analyzing and studying MCER issues is significant to…

Artificial Intelligence · Computer Science 2025-11-14 Yuntao Shou , Tao Meng , Wei Ai , Fangze Fu , Nan Yin , Keqin Li

Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redundant cross-modal…

Sound · Computer Science 2026-04-17 Chengling Guo , Yuntao Shou , Tao Meng , Wei Ai , Yun Tan , Keqin Li

The rapid expansion of the electric vehicle (EV) industry has highlighted the importance of user feedback in improving product design and charging infrastructure. Traditional sentiment analysis methods often oversimplify the complexity of…

Artificial Intelligence · Computer Science 2024-12-06 Shuhao Chen , Chengyi Tu

Employing voice-based emotion recognition function in artificial intelligence (AI) product will improve the user experience. Most of researches that have been done only focus on the speech collected under controlled conditions. The…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-06 Fei Tao , Gang Liu , Qingen Zhao

Automatic speech emotion recognition (SER) by a computer is a critical component for more natural human-machine interaction. As in human-human interaction, the capability to perceive emotion correctly is essential to take further steps in a…

Sound · Computer Science 2022-10-27 Bagus Tris Atmaja , Masato Akagi

Emotion detection from text seeks to identify an individual's emotional or mental state - positive, negative, or neutral - based on linguistic cues. While significant progress has been made for English and other high-resource languages,…

Computation and Language · Computer Science 2025-11-11 Abdullah Al Maruf , Aditi Golder , Zakaria Masud Jiyad , Abdullah Al Numan , Tarannum Shaila Zaman

Emotion recognition has a wide range of applications in human-computer interaction, marketing, healthcare, and other fields. In recent years, the development of deep learning technology has provided new methods for emotion recognition.…

Computation and Language · Computer Science 2025-01-28 Junwei Feng , Xueyan Fan

In this study, the Multivariate Empirical Mode Decomposition (MEMD) is applied to multichannel EEG to obtain scale-aligned intrinsic mode functions (IMFs) as input features for emotion detection. The IMFs capture local signal variation…

Signal Processing · Electrical Eng. & Systems 2022-06-03 Monira Islam , Tan Lee

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Haobin Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a…

Sound · Computer Science 2026-03-06 Cong Wang , Yizhong Geng , Yuhua Wen , Qifei Li , Yingming Gao , Ruimin Wang , Chunfeng Wang , Hao Li , Ya Li , Wei Chen

There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited number of emotion types and have achieved unsatisfactory…

Sound · Computer Science 2023-06-02 Haobin Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao