中文
相关论文

相关论文: VISTANet: VIsual Spoken Textual Additive Net for I…

200 篇论文

This paper investigates the influence of different acoustic features, audio-events based features and automatic speech translation based lexical features in complex emotion recognition such as curiosity. Pretrained networks, namely,…

声音 · 计算机科学 2018-11-05 Bhalaji Nagarajan , V Ramana Murthy Oruganti

Inspired from the assets of handcrafted and deep learning approaches, we proposed a RARITYNet: RARITY guided affective emotion learning framework to learn the appearance features and identify the emotion class of facial expressions. The…

计算机视觉与模式识别 · 计算机科学 2022-05-19 Monu Verma , Santosh Kumar Vipparthi

We introduce a novel multimodal emotion recognition dataset that enhances the precision of Valence-Arousal Model while accounting for individual differences. This dataset includes electroencephalography (EEG), electrocardiography (ECG), and…

人机交互 · 计算机科学 2025-03-24 Xin Huang , Shiyao Zhu , Ziyu Wang , Yaping He , Hao Jin , Zhengkui Liu

Visual emotion analysis (VEA) has attracted great attention recently, due to the increasing tendency of expressing and understanding emotions through images on social networks. Different from traditional vision tasks, VEA is inherently more…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Jingyuan Yang , Jie Li , Xiumei Wang , Yuxuan Ding , Xinbo Gao

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great capacity of transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Baoshuo Kan , Teng Wang , Wenpeng Lu , Xiantong Zhen , Weili Guan , Feng Zheng

Facial emotion recognition has been typically cast as a single-label classification problem of one out of six prototypical emotions. However, that is an oversimplification that is unsuitable for representing the multifaceted spectrum of…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Joao Baptista Cardia Neto , Claudio Ferrari , Stefano Berretti

Humans are able to comprehend information from multiple domains for e.g. speech, text and visual. With advancement of deep learning technology there has been significant improvement of speech recognition. Recognizing emotion from speech is…

音频与语音处理 · 电气工程与系统科学 2020-06-16 Mandeep Singh , Yuan Fang

In this paper, we introduce the semantic knowledge of medical images from their diagnostic reports to provide an inspirational network training and an interpretable prediction mechanism with our proposed novel multimodal neural network,…

计算机视觉与模式识别 · 计算机科学 2017-08-11 Zizhao Zhang , Pingjun Chen , Manish Sapkota , Lin Yang

Human emotion is expressed in many communication modalities and media formats and so their computational study is equally diversified into natural language processing, audio signal analysis, computer vision, etc. Similarly, the large…

机器学习 · 计算机科学 2023-08-16 Sven Buechel , Udo Hahn

Robust speech emotion recognition relies on the quality of the speech features. We present speech features enhancement strategy that improves speech emotion recognition. We used the INTERSPEECH 2010 challenge feature-set. We identified…

信号处理 · 电气工程与系统科学 2022-08-22 Sofia Kanwal , Sohail Asghar , Hazrat Ali

Multimodal emotion recognition is a challenging task in emotion computing as it is quite difficult to extract discriminative features to identify the subtle differences in human emotions with abstract concept and multiple expressions.…

声音 · 计算机科学 2021-11-18 Hengshun Zhou , Jun Du , Yuanyuan Zhang , Qing Wang , Qing-Feng Liu , Chin-Hui Lee

Automated Facial Expression Recognition (FER) is challenging due to intra-class variations and inter-class similarities. FER can be especially difficult when facial expressions reflect a mixture of various emotions (aka compound…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Ali Pourramezan Fard , Mohammad Mehdi Hosseini , Timothy D. Sweeny , Mohammad H. Mahoor

We developed a novel, interpretable multimodal classification method to identify symptoms of mood disorders viz. depression, anxiety and anhedonia using audio, video and text collected from a smartphone application. We used CNN-based…

For a long time, images have proved perfect at both storing and conveying rich semantics, especially human emotions. A lot of research has been conducted to provide machines with the ability to recognize emotions in photos of people.…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Quoc-Bao Ninh , Hai-Chan Nguyen , Triet Huynh , Trung-Nghia Le

Recent progress in semantic segmentation is driven by deep Convolutional Neural Networks and large-scale labeled image datasets. However, data labeling for pixel-wise segmentation is tedious and costly. Moreover, a trained model can only…

计算机视觉与模式识别 · 计算机科学 2019-03-07 Chi Zhang , Guosheng Lin , Fayao Liu , Rui Yao , Chunhua Shen

The explosive increase of multimodal data makes a great demand in many cross-modal applications that follow the strict prior related assumption. Thus researchers study the definition of cross-modal correlation category and construct various…

计算机视觉与模式识别 · 计算机科学 2021-09-03 Nan Xu , Junyan Wang , Yuan Tian , Ruike Zhang , Wenji Mao

While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla…

音频与语音处理 · 电气工程与系统科学 2023-02-22 Li-Wei Chen , Alexander Rudnicky

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

声音 · 计算机科学 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Messages in human conversations inherently convey emotions. The task of detecting emotions in textual conversations leads to a wide range of applications such as opinion mining in social networks. However, enabling machines to analyze…

计算与语言 · 计算机科学 2019-10-02 Peixiang Zhong , Di Wang , Chunyan Miao

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Tai Vu