English
Related papers

Related papers: GM-TCNet: Gated Multi-scale Temporal Convolutional…

200 papers

Emotion recognition is a crucial task for human conversation understanding. It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions. As a typical solution, the global- and the local…

Computation and Language · Computer Science 2024-01-31 Cam-Van Thi Nguyen , Anh-Tuan Mai , The-Son Le , Hai-Dang Kieu , Duc-Trong Le

In this work, we address the problem of finegrained traceback of emotional and manipulation characteristics from synthetically manipulated speech. We hypothesize that combining semantic-prosodic cues captured by Speech Foundation Models…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-17 Girish , Mohd Mujtaba Akhtar , Farhan Sheth , Muskaan Singh

Speech emotion recognition (SER) is pivotal for enhancing human-machine interactions. This paper introduces "EmoHRNet", a novel adaptation of High-Resolution Networks (HRNet) tailored for SER. The HRNet structure is designed to maintain…

Sound · Computer Science 2025-10-08 Akshay Muppidi , Martin Radfar

The task of multi-modal emotion recognition in conversation (MERC) aims to analyze the genuine emotional state of each utterance based on the multi-modal information in the conversation, which is crucial for conversation understanding.…

Machine Learning · Computer Science 2024-09-04 Yuntao Shou , Wei Ai , Jiayi Du , Tao Meng , Haiyan Liu , Nan Yin

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP)…

Sound · Computer Science 2025-07-08 Jiacheng Shi , Yanfu Zhang , Ye Gao

Multimodal machine learning is an emerging area of research, which has received a great deal of scholarly attention in recent years. Up to now, there are few studies on multimodal Emotion Recognition in Conversation (ERC). Since Graph…

Multimedia · Computer Science 2023-12-05 Jiang Li , Xiaoping Wang , Guoqing Lv , Zhigang Zeng

Recent advances in text-to-speech (TTS) have enabled natural speech synthesis, but fine-grained, time-varying emotion control remains challenging. Existing methods often allow only utterance-level control and require full model fine-tuning…

Sound · Computer Science 2025-07-08 Jaeseok Jeong , Yuna Lee , Mingi Kwon , Youngjung Uh

Speech Emotion Recognition (SER) is crucial for human-computer interaction but still remains a challenging problem because of two major obstacles: data scarcity and imbalance. Many datasets for SER are substantially imbalanced, where data…

Sound · Computer Science 2022-08-11 Shijun Wang , Hamed Hemati , Jón Guðnason , Damian Borth

Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with linguistic cues,…

Sound · Computer Science 2023-07-03 Anna Ollerenshaw , Md Asif Jalal , Rosanna Milner , Thomas Hain

Speech emotion recognition (SER) is to study the formation and change of speaker's emotional state from the speech signal perspective, so as to make the interaction between human and computer more intelligent. SER is a challenging task that…

Sound · Computer Science 2017-08-01 Yafeng Niu , Dongsheng Zou , Yadong Niu , Zhongshi He , Hua Tan

The research on human emotion under multimedia stimulation based on physiological signals is an emerging field, and important progress has been achieved for emotion recognition based on multi-modal signals. However, it is challenging to…

Machine Learning · Computer Science 2021-08-10 Ziyu Jia , Youfang Lin , Jing Wang , Zhiyang Feng , Xiangheng Xie , Caijie Chen

Learning expressive representation is crucial in deep learning. In speech emotion recognition (SER), vacuum regions or noises in the speech interfere with expressive representation learning. However, traditional RNN-based models are…

Sound · Computer Science 2022-08-23 Junghun Kim , Jihie Kim

Automatic affect recognition is a challenging task due to the various modalities emotions can be expressed with. Applications can be found in many domains including multimedia retrieval and human computer interaction. In recent years, deep…

Computer Vision and Pattern Recognition · Computer Science 2018-02-14 Panagiotis Tzirakis , George Trigeorgis , Mihalis A. Nicolaou , Björn Schuller , Stefanos Zafeiriou

This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We collect the personality annotation for the IEMOCAP dataset,…

Sound · Computer Science 2025-12-01 Yuan Gao , Hao Shi , Yahui Fu , Chenhui Chu , Tatsuya Kawahara

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a challenging problem,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Lucas Ueda , João Lima , Leonardo Marques , Paula Costa

Using mel-spectrograms over conventional MFCCs features, we assess the abilities of convolutional neural networks to accurately recognize and classify emotions from speech data. We introduce FSER, a speech emotion recognition model trained…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-17 Bonaventure F. P. Dossou , Yeno K. S. Gbenou

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

Sound · Computer Science 2026-05-11 Yassin Terraf , Youssef Iraqi

Knowledge of users' emotion states helps improve human-computer interaction. In this work, we presented EmoNet, an emotion detector of Chinese daily dialogues based on deep convolutional neural networks. In order to maintain the original…

Computation and Language · Computer Science 2017-10-04 Jialiang Zhao , Qi Gao

Speech emotion recognition (SER) has received a great deal of attention in recent years in the context of spontaneous conversations. While there have been notable results on datasets like the well known corpus of naturalistic dyadic…

Computation and Language · Computer Science 2024-01-02 Alex-Răzvan Ispas , Théo Deschamps-Berger , Laurence Devillers

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson