中文
相关论文

相关论文: Speech Emotion Recognition with Global-Aware Fusio…

200 篇论文

Speech Emotion Recognition (SER) is crucial in human-machine interactions. Mainstream approaches utilize Convolutional Neural Networks or Recurrent Neural Networks to learn local energy feature representations of speech segments from speech…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Xiaoyu Tang , Yixin Lin , Ting Dang , Yuanfang Zhang , Jintao Cheng

Speech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lacks comprehensive spatiotemporal feature learning. This paper…

声音 · 计算机科学 2023-12-29 Mengbo Li , Yuanzhong Zheng , Dichucheng Li , Yulun Wu , Yaoxuan Wang , Haojun Fei

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a…

声音 · 计算机科学 2026-03-06 Cong Wang , Yizhong Geng , Yuhua Wen , Qifei Li , Yingming Gao , Ruimin Wang , Chunfeng Wang , Hao Li , Ya Li , Wei Chen

A graph neural network (GNN) for image understanding based on multiple cues is proposed in this paper. Compared to traditional feature and decision fusion approaches that neglect the fact that features can interact and exchange information,…

计算机视觉与模式识别 · 计算机科学 2020-03-02 Xin Guo , Luisa F. Polania , Bin Zhu , Charles Boncelet , Kenneth E. Barner

Owing to the recent developments in Generative Artificial Intelligence (GenAI) and Large Language Models (LLM), conversational agents are becoming increasingly popular and accepted. They provide a human touch by interacting in ways familiar…

计算与语言 · 计算机科学 2023-10-31 Fathima Abdul Rahman , Guang Lu

Speech emotion recognition (SER) is an essential part of human-computer interaction. In this paper, we propose an SER network based on a Graph Isomorphism Network with Weighted Multiple Aggregators (WMA-GIN), which can effectively handle…

音频与语音处理 · 电气工程与系统科学 2022-11-02 Ying Hu , Yuwu Tang , Hao Huang , Liang He

Multimodal emotion recognition in conversations (MERC) aims to identify and understand the emotions expressed by speakers during utterance interaction from multiple modalities (e.g., text, audio, images, etc.). Existing studies have shown…

人工智能 · 计算机科学 2026-03-25 Tao Meng , Weilun Tang , Yuntao Shou , Yilong Tan , Jun Zhou , Wei Ai , Keqin Li

Transformers have proved to be very effective for visual recognition tasks. In particular, vision transformers construct compressed global representations through self-attention and learnable class tokens. Multi-resolution transformers have…

计算机视觉与模式识别 · 计算机科学 2022-12-16 Loic Themyr , Clement Rambour , Nicolas Thome , Toby Collins , Alexandre Hostettler

Speech emotion recognition (SER) is crucial for enhancing affective computing and enriching the domain of human-computer interaction. However, the main challenge in SER lies in selecting relevant feature representations from speech signals…

声音 · 计算机科学 2024-12-16 Niloy Kumar Kundu , Sarah Kobir , Md. Rayhan Ahmed , Tahmina Aktar , Niloya Roy

In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attention granularity. In…

声音 · 计算机科学 2021-02-04 Mingke Xu , Fan Zhang , Xiaodong Cui , Wei Zhang

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025…

The performance of speech emotion recognition (SER) is limited by the insufficient emotion information in unimodal systems and the feature alignment difficulties in multimodal systems. Recently, multimodal large language models (MLLMs) have…

声音 · 计算机科学 2025-09-22 Yiqing Yang , Man-Wai Mak

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Xuechen Wang , Shiwan Zhao , Haoqin Sun , Hui Wang , Jiaming Zhou , Yong Qin

Speech Emotion Recognition (SER) aims to help the machine to understand human's subjective emotion from only audio information. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this…

声音 · 计算机科学 2022-03-30 Heqing Zou , Yuke Si , Chen Chen , Deepu Rajan , Eng Siong Chng

Conversational emotion recognition (CER) is an important research topic in human-computer interactions. {Although recent advancements in transformer-based cross-modal fusion methods have shown promise in CER tasks, they tend to overlook the…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Yuntao Shou , Huan Liu , Xiangyong Cao , Deyu Meng , Bo Dong

In recent years, powered by the learned discriminative representation via graph neural network (GNN) models, deep graph matching methods have made great progresses in the task of matching semantic features. However, these methods usually…

计算机视觉与模式识别 · 计算机科学 2021-11-18 He Liu , Tao Wang , Yidong Li , Congyan Lang , Yi Jin , Haibin Ling

Emotion recognition from speech is a challenging task. Re-cent advances in deep learning have led bi-directional recur-rent neural network (Bi-RNN) and attention mechanism as astandard method for speech emotion recognition, extractingand…

声音 · 计算机科学 2021-06-09 Zixuan Peng , Yu Lu , Shengfeng Pan , Yunfeng Liu

Since Multimodal Emotion Recognition in Conversation (MERC) can be applied to public opinion monitoring, intelligent dialogue robots, and other fields, it has received extensive research attention in recent years. Unlike traditional…

机器学习 · 计算机科学 2024-07-25 Tao Meng , Fuchen Zhang , Yuntao Shou , Hongen Shao , Wei Ai , Keqin Li

Continuous dimensional speech emotion recognition captures affective variation along valence, arousal, and dominance, providing finer-grained representations than categorical approaches. Yet most multimodal methods rely solely on global…

声音 · 计算机科学 2026-01-27 Haoxun Li , Yuqing Sun , Hanlei Shi , Yu Liu , Leyuan Qu , Taihao Li

This study investigates fine-tuning self-supervised learn ing (SSL) models using multi-task learning (MTL) to enhance speech emotion recognition (SER). The framework simultane ously handles four related tasks: emotion recognition, gender…

声音 · 计算机科学 2025-08-26 Honghong Wang , Jing Deng , Fanqin Meng , Rong Zheng
‹ 上一页 1 2 3 10 下一页 ›