中文
相关论文

相关论文: WavFusion: Towards wav2vec 2.0 Multimodal Speech E…

200 篇论文

Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed…

机器学习 · 计算机科学 2022-05-03 Ahmed Abdou , Ekta Sood , Philipp Müller , Andreas Bulling

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in…

声音 · 计算机科学 2025-12-23 Sudip Chakrabarty , Pappu Bishwas , Rajdeep Chatterjee

Speech emotion recognition is a challenging task and an important step towards more natural human-machine interaction. We show that pre-trained language models can be fine-tuned for text emotion recognition, achieving an accuracy of 69.5%…

音频与语音处理 · 电气工程与系统科学 2019-12-06 Verena Heusser , Niklas Freymuth , Stefan Constantin , Alex Waibel

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

计算与语言 · 计算机科学 2021-01-27 Gaurav Sahu , Olga Vechtomova

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

声音 · 计算机科学 2023-09-29 R. Gnana Praveen , Jahangir Alam

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Tai Vu

In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech…

音频与语音处理 · 电气工程与系统科学 2025-09-26 Jing-Tong Tzeng , Bo-Hao Su , Ya-Tse Wu , Hsing-Hang Chou , Chi-Chun Lee

Due to its ability to accurately predict emotional state using multimodal features, audiovisual emotion recognition has recently gained more interest from researchers. This paper proposes two methods to predict emotional attributes from…

音频与语音处理 · 电气工程与系统科学 2022-07-22 Bagus Tris Atmaja , Masato Akagi

A multi-modal emotion recognition method was established by combining two-channel convolutional neural network with ring network. This method can extract emotional information effectively and improve learning efficiency. The words were…

人工智能 · 计算机科学 2023-11-21 Jiazhen Wang

Emotion Recognition (ER) is the process of analyzing and identifying human emotions from sensing data. Currently, the field heavily relies on facial expression recognition (FER) because visual channel conveys rich emotional cues. However,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Kejun Liu , Yuanyuan Liu , Lin Wei , Chang Tang , Yibing Zhan , Zijing Chen , Zhe Chen

Emotion recognition has become an important field of research in Human Computer Interactions as we improve upon the techniques for modelling the various aspects of behaviour. With the advancement of technology our understanding of emotions…

人工智能 · 计算机科学 2019-11-11 Samarth Tripathi , Sarthak Tripathi , Homayoon Beigi

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

计算机视觉与模式识别 · 计算机科学 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature…

声音 · 计算机科学 2024-09-10 Pujin Shi , Fei Gao

Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker…

声音 · 计算机科学 2025-11-19 Xiao Li , Kotaro Funakoshi , Manabu Okumura

Multimodal emotion recognition is an important research topic in artificial intelligence. Over the past few decades, researchers have made remarkable progress by increasing the dataset size and building more effective algorithms. However,…

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for…

计算与语言 · 计算机科学 2019-04-24 Pan Zhou , Wenwen Yang , Wei Chen , Yanfeng Wang , Jia Jia

We used two multimodal models for continuous valence-arousal recognition using visual, audio, and linguistic information. The first model is the same as we used in ABAW2 and ABAW3, which employs the leader-follower attention. The second…

多媒体 · 计算机科学 2023-04-18 Su Zhang , Ziyuan Zhao , Cuntai Guan

Automatic emotion recognition is an active research topic with wide range of applications. Due to the high manual annotation cost and inevitable label ambiguity, the development of emotion recognition dataset is limited in both scale and…

音频与语音处理 · 电气工程与系统科学 2020-09-08 Jingjun Liang , Ruichen Li , Qin Jin

In the domain of human-computer interaction, accurately recognizing and interpreting human emotions is crucial yet challenging due to the complexity and subtlety of emotional expressions. This study explores the potential for detecting a…

多媒体 · 计算机科学 2025-05-13 Jiehui Jia , Huan Zhang , Jinhua Liang