English
Related papers

Related papers: Advancing Audio Emotion and Intent Recognition wit…

200 papers

Automatic speech emotion recognition (SER) by a computer is a critical component for more natural human-machine interaction. As in human-human interaction, the capability to perceive emotion correctly is essential to take further steps in a…

Sound · Computer Science 2022-10-27 Bagus Tris Atmaja , Masato Akagi

The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal…

Sound · Computer Science 2025-12-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

Multimodal speech emotion recognition (SER) has emerged as pivotal for improving human-machine interaction. Researchers are increasingly leveraging both speech and textual information obtained through automatic speech recognition (ASR) to…

Human-Computer Interaction · Computer Science 2025-09-24 Jiajun He , Xiaohan Shi , Cheng-Hung Hu , Jinyi Mi , Xingfeng Li , Tomoki Toda

Analysis of human affect plays a vital role in human-computer interaction (HCI) systems. Due to the difficulty in capturing large amounts of real-life data, most of the current methods have mainly focused on controlled environments, which…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Jun Yu , Zhongpeng Cai , Peng He , Guocheng Xie , Qiang Ling

Human emotion can be presented in different modes i.e., audio, video, and text. However, the contribution of each mode in exhibiting each emotion is not uniform. Furthermore, the availability of complete mode-specific details may not always…

Artificial Intelligence · Computer Science 2024-02-20 Naresh Kumar Devulapally , Sidharth Anand , Sreyasee Das Bhattacharjee , Junsong Yuan

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during…

Computation and Language · Computer Science 2025-08-29 Weiting Tan , Jiachen Lian , Hirofumi Inaguma , Paden Tomasello , Philipp Koehn , Xutai Ma

Despite the success of deep learning in speech recognition, multi-dialect speech recognition remains a difficult problem. Although dialect-specific acoustic models are known to perform well in general, they are not easy to maintain when…

Machine Learning · Computer Science 2022-05-09 Sanghyun Yoo , Inchul Song , Yoshua Bengio

In this paper, we present our advanced solutions to the two sub-challenges of Affective Behavior Analysis in the wild (ABAW) 2023: the Emotional Reaction Intensity (ERI) Estimation Challenge and Expression (Expr) Classification Challenge.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Jia Li , Yin Chen , Xuesong Zhang , Jiantao Nie , Ziqiang Li , Yangchen Yu , Yan Zhang , Richang Hong , Meng Wang

Despite the abundance of current researches working on the sentiment analysis from videos and audios, finding the best model that gives the highest accuracy rate is still considered a challenge for researchers in this field. The main…

Sound · Computer Science 2024-12-13 Antonio Fernandez , Suzan Awinat

As artificial intelligence (AI) systems become increasingly embedded in our daily life, the ability to recognize and adapt to human emotions is essential for effective human-computer interaction. Facial expression recognition (FER) provides…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Thibault Geoffroy , Myriam Maumy , Lionel Prevost

This paper reports the analysis of audio and visual features in predicting the continuous emotion dimensions under the seventh Audio/Visual Emotion Challenge (AVEC 2017), which was done as part of a B.Tech. 2nd year internship project. For…

Computer Vision and Pattern Recognition · Computer Science 2017-10-25 Narotam Singh , Nittin Singh , Abhinav Dhall

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a…

Sound · Computer Science 2025-04-02 Jiachen Luo , Huy Phan , Lin Wang , Joshua Reiss

This paper presents a detailed system description of our entry for the WASSA 2024 Task 2, focused on cross-lingual emotion detection. We utilized a combination of large language models (LLMs) and their ensembles to effectively understand…

Computation and Language · Computer Science 2024-10-22 Ram Mohan Rao Kadiyala

Acoustic-to-word (A2W) models that allow direct mapping from acoustic signals to word sequences are an appealing approach to end-to-end automatic speech recognition due to their simplicity. However, prior works have shown that modelling A2W…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-17 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a single dominant signal. We propose a rank-aware multi-encoder framework that selectively…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Junghyun Lee , Hyunseo Kim , Hanna Jang , Junhyug Noh

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

We present our submission to the Hume-ABAW10 Emotional Mimicry Intensity (EMI) Challenge, which aims to predict six continuous emotion intensity dimensions: Admiration, Amusement, Determination, Empathic Pain, Excitement, and Joy, from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Dinithi Dissanayake , Shaveen Silva , Ovindu Atukorala , Prasanth Sasikumar , Suranga Nanayakkara

Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to effectively adapt these representations for speech emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-28 Sundararajan Srinivasan , Zhaocheng Huang , Katrin Kirchhoff

Transformer models have significantly advanced the field of emotion recognition. However, there are still open challenges when exploring open-ended queries for Large Language Models (LLMs). Although current models offer good results,…