English
Related papers

Related papers: DNN-HMM based Speaker Adaptive Emotion Recognition…

200 papers

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

Speech is a natural form of communication for human beings, and computers with the ability to understand speech and speak with a human voice are expected to contribute to the development of more natural man-machine interfaces. Computers…

Sound · Computer Science 2013-05-15 Neema Mishra , Urmila Shrawankar , V M Thakare

We propose a learnable mel-frequency cepstral coefficient (MFCC) frontend architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features…

Sound · Computer Science 2021-02-23 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

In this paper, we propose a novel family of windowing technique to compute Mel Frequency Cepstral Coefficient (MFCC) for automatic speaker recognition from speech. The proposed method is based on fundamental property of discrete time…

Computer Vision and Pattern Recognition · Computer Science 2015-06-05 Md. Sahidullah , Goutam Saha

In this paper, we propose to use deep 3-dimensional convolutional networks (3D CNNs) in order to address the challenge of modelling spectro-temporal dynamics for speech emotion recognition (SER). Compared to a hybrid of Convolutional Neural…

Computation and Language · Computer Science 2017-08-18 Jaebok Kim , Khiet P. Truong , Gwenn Englebienne , Vanessa Evers

Traditionally, in paralinguistic analysis for emotion detection from speech, emotions have been identified with discrete or dimensional (continuous-valued) labels. Accordingly, models that have been proposed for emotion detection use one or…

Sound · Computer Science 2022-11-01 Roshan Sharma , Hira Dhamyal , Bhiksha Raj , Rita Singh

In this work, we conducted an empirical comparative study of the performance of text-independent speaker verification in emotional and stressful environments. This work combined deep models with shallow architecture, which resulted in novel…

Sound · Computer Science 2021-12-28 Ismail Shahin , Ali Bou Nassif , Nawel Nemmour , Ashraf Elnagar , Adi Alhudhaif , Kemal Polat

The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English speakers of…

Sound · Computer Science 2022-06-16 Pierre Berjon , Rajib Sharma , Avishek Nag , Soumyabrata Dev

Emotions recognition is commonly employed for health assessment. However, the typical metric for evaluation in therapy is based on patient-doctor appraisal. This process can fall into the issue of subjectivity, while also requiring…

Human-Computer Interaction · Computer Science 2021-01-21 Jumana Almahmoud , Kruthika Kikkeri

Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-25 Luz Martinez-Lucas , Pravin Mote , Abinay Reddy Naini , Mohammed Abdelwahab , Carlos Busso

This research is dedicated to improving text-independent Emirati-accented speaker identification performance in stressful talking conditions using three distinct classifiers: First-Order Hidden Markov Models (HMM1s), Second-Order Hidden…

Sound · Computer Science 2019-10-30 Ismail Shahin , Ali Bou Nassif

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text…

Computation and Language · Computer Science 2018-02-26 Yue Gu , Shuhong Chen , Ivan Marsic

Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time-frequency pattern of speech signal for speech emotion recognition (SER). \textcolor{black}{Generally, different emotions correspond…

Sound · Computer Science 2022-10-25 Cheng Lu , Wenming Zheng , Hailun Lian , Yuan Zong , Chuangao Tang , Sunan Li , Yan Zhao

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

Computation and Language · Computer Science 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel…

Sound · Computer Science 2025-10-03 Jianing Yang , Sheng Li , Takahiro Shinozaki , Yuki Saito , Hiroshi Saruwatari

Emotion Recognition in Conversations (ERC) is a key step towards successful human-machine interaction. While the field has seen tremendous advancement in the last few years, new applications and implementation scenarios present novel…

Computation and Language · Computer Science 2024-10-22 Patrícia Pereira , Helena Moniz , Joao Paulo Carvalho

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in speech emotion…

Computation and Language · Computer Science 2014-06-25 Imen Trabelsi , Dorra Ben Ayed , Noureddine Ellouze

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While…

Sound · Computer Science 2026-02-13 Chung-Soo Ahn , Rajib Rana , Sunil Sivadas , Carlos Busso , Jagath C. Rajapakse

Despite strong recent progress in Emotion Recognition in Conversation (ERC), two gaps remain: we lack clear understanding of which modeling choices materially affect performance, and we have limited linguistic analysis linking recognition…

Computation and Language · Computer Science 2026-02-10 Cheonkam Jeong , Adeline Nyamathi