English
Related papers

Related papers: VISTANet: VIsual Spoken Textual Additive Net for I…

200 papers

Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective…

Sound · Computer Science 2020-07-29 Siddique Latif , Rajib Rana , Junaid Qadir , Julien Epps

We propose a framework for multimodal sentiment analysis and emotion recognition using convolutional neural network-based feature extraction from text and visual modalities. We obtain a performance improvement of 10% over the state of the…

Multimedia · Computer Science 2017-08-01 Erik Cambria , Devamanyu Hazarika , Soujanya Poria , Amir Hussain , R. B. V. Subramaanyam

In recent times, there has been significant interest in the machine recognition of human emotions, due to the suite of applications to which this knowledge can be applied. A number of different modalities, such as speech or facial…

Human-Computer Interaction · Computer Science 2018-03-06 Jonny O'Dwyer , Ronan Flynn , Niall Murray

In this work, we propose different variants of the self-attention based network for emotion prediction from movies, which we call AttendAffectNet. We take both audio and video into account and incorporate the relation among multiple…

Sound · Computer Science 2021-10-19 Ha Thi Phuong Thao , Balamurali B. T. , Dorien Herremans , Gemma Roig

Leveraging the visual modality effectively for Neural Machine Translation (NMT) remains an open problem in computational linguistics. Recently, Caglayan et al. posit that the observed gains are limited mainly due to the very simple, short,…

Computation and Language · Computer Science 2019-10-08 Vikas Raunak , Sang Keun Choe , Quanyang Lu , Yi Xu , Florian Metze

Sentiment analysis and emotion recognition in videos are challenging tasks, given the diversity and complexity of the information conveyed in different modalities. Developing a highly competent framework that effectively addresses the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Prasad Chaudhari , Aman Kumar , Chandravardhan Singh Raghaw , Mohammad Zia Ur Rehman , Nagendra Kumar

Classifying group-level emotions is a challenging task due to complexity of video, in which not only visual, but also audio information should be taken into consideration. Existing works on multimodal emotion recognition are using bulky…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Lev Evtodienko

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the…

Technological advancements in web platforms allow people to express and share emotions towards textual write-ups written and shared by others. This brings about different interesting domains for analysis; emotion expressed by the writer and…

Computation and Language · Computer Science 2025-03-12 Anoop Kadan , Deepak P. , Manjary P. Gangan , Savitha Sam Abraham , Lajish V. L

In this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and valence levels. The proposed approach employs a pre-trained…

Machine Learning · Computer Science 2019-06-26 Juan D. S. Ortega , Patrick Cardinal , Alessandro L. Koerich

Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Guanyu Hu , Dimitrios Kollias , Xinyu Yang

ERIT is a novel multimodal dataset designed to facilitate research in a lightweight multimodal fusion. It contains text and image data collected from videos of elderly individuals reacting to various situations, as well as seven emotion…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Rita Frieske , Bertram E. Shi

In this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Licheng Yu , Zhe Lin , Xiaohui Shen , Jimei Yang , Xin Lu , Mohit Bansal , Tamara L. Berg

Emotions widely affect human decision-making. This fact is taken into account by affective computing with the goal of tailoring decision support to the emotional states of individuals. However, the accurate recognition of emotions within…

Computation and Language · Computer Science 2018-11-14 Bernhard Kratzwald , Suzana Ilic , Mathias Kraus , Stefan Feuerriegel , Helmut Prendinger

In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and…

Sound · Computer Science 2023-02-20 Yan Zhao , Jincen Wang , Yuan Zong , Wenming Zheng , Hailun Lian , Li Zhao

Vision-language models (VLMs) show promise as tools for inferring affect from visual stimuli at scale; it is not yet clear how closely their outputs align with human affective ratings. We benchmarked nine VLMs, ranging from state-of-the-art…

Recent advancements in sequence prediction have significantly improved the accuracy of video data interpretation; however, existing models often overlook the potential of attention-based mechanisms for next-frame prediction. This study…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Yiqiao Yin

Depression is a severe global mental health issue that impairs daily functioning and overall quality of life. Although recent audio-visual approaches have improved automatic depression detection, methods that ignore emotional cues often…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Chenglizhao Chen , Boze Li , Mengke Song , Dehao Feng , Xinyu Liu , Shanchen Pang , Jufeng Yang , Hui Yu

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

Inspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semantic attributes for…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Liang Sun , Bing Li , Chunfeng Yuan , Zhengjun Zha , Weiming Hu