English
Related papers

Related papers: Multi-Granularity Network with Modal Attention for…

200 papers

Multimodal Emotion Recognition (MER) aims to perceive human emotions through three modes: language, vision, and audio. Previous methods primarily focused on modal fusion without adequately addressing significant distributional differences…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Jichao Zhu , Jun Yu

Current state-of-the-art neural machine translation (NMT) uses a deep multi-head self-attention network with no explicit phrase information. However, prior work on statistical machine translation has shown that extending the basic…

Computation and Language · Computer Science 2019-09-06 Jie Hao , Xing Wang , Shuming Shi , Jinfeng Zhang , Zhaopeng Tu

The audio-video based multimodal emotion recognition has attracted a lot of attention due to its robust performance. Most of the existing methods focus on proposing different cross-modal fusion strategies. However, these strategies…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Ziwang Fu , Feng Liu , Hanyang Wang , Jiayin Qi , Xiangling Fu , Aimin Zhou , Zhibin Li

Visual attention mechanisms are a key component of neural network models for computer vision. By focusing on a discrete set of objects or image regions, these mechanisms identify the most relevant features and use them to build more…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 António Farinhas , André F. T. Martins , Pedro M. Q. Aguiar

Multimodal sentiment analysis (MSA) leverages information fusion from diverse modalities (e.g., text, audio, visual) to enhance sentiment prediction. However, simple fusion techniques often fail to account for variations in modality…

Machine Learning · Computer Science 2025-10-03 Han Wu , Yanming Sun , Yunhe Yang , Derek F. Wong

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

Sound · Computer Science 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Recently, emotion recognition based on physiological signals has emerged as a field with intensive research. The utilization of multi-modal, multi-channel physiological signals has significantly improved the performance of emotion…

Multimedia · Computer Science 2023-08-22 Xinda Li

Emotional expressions are the behaviors that communicate our emotional state or attitude to others. They are expressed through verbal and non-verbal communication. Complex human behavior can be understood by studying physical features from…

Computer Vision and Pattern Recognition · Computer Science 2021-09-15 Liam Schoneveld , Alice Othmani , Hazem Abdelkawy

Nowadays, numerous online platforms can be described as multi-modal heterogeneous networks (MMHNs), such as Douban's movie networks and Amazon's product review networks. Accurately categorizing nodes within these networks is crucial for…

Machine Learning · Computer Science 2025-06-23 Jiafan Li , Jiaqi Zhu , Liang Chang , Yilin Li , Miaomiao Li , Yang Wang , Hongan Wang

Multimodal brain networks characterize complex connectivities among different brain regions from both structural and functional aspects and provide a new means for mental disease analysis. Recently, Graph Neural Networks (GNNs) have become…

Neurons and Cognition · Quantitative Biology 2022-05-25 Yanqiao Zhu , Hejie Cui , Lifang He , Lichao Sun , Carl Yang

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Haotian Wang , Yuxuan Xi , Hang Chen , Jun Du , Yan Song , Qing Wang , Hengshun Zhou , Chenxi Wang , Jiefeng Ma , Pengfei Hu , Ya Jiang , Shi Cheng , Jie Zhang , Yuzhe Weng

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Junhyeong Byeon , Jeongyeol Kim , Sejoon Lim

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial attention mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2017-03-13 Chiori Hori , Takaaki Hori , Teng-Yok Lee , Kazuhiro Sumi , John R. Hershey , Tim K. Marks

In this work, a kernel attention module is presented for the task of EEG-based emotion classification with neural networks. The proposed module utilizes a self-attention mechanism by performing a kernel trick, demanding significantly fewer…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Dongyang Kuang , Craig Michoski

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

Computation and Language · Computer Science 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu

In this work, we explore the impact of visual modality in addition to speech and text for improving the accuracy of the emotion detection system. The traditional approaches tackle this task by fusing the knowledge from the various…

Machine Learning · Computer Science 2020-04-24 Seunghyun Yoon , Subhadeep Dey , Hwanhee Lee , Kyomin Jung

Multi-modal entity alignment (MMEA) aims to discover identical entities across different knowledge graphs (KGs) whose entities are associated with relevant images. However, current MMEA algorithms rely on KG-level modality fusion strategies…

Artificial Intelligence · Computer Science 2023-08-01 Zhuo Chen , Jiaoyan Chen , Wen Zhang , Lingbing Guo , Yin Fang , Yufeng Huang , Yichi Zhang , Yuxia Geng , Jeff Z. Pan , Wenting Song , Huajun Chen

Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Weidong Chen , Dexiang Hong , Yuankai Qi , Zhenjun Han , Shuhui Wang , Laiyun Qing , Qingming Huang , Guorong Li

Emotion recognition based on Electroencephalography (EEG) has gained significant attention and diversified development in fields such as neural signal processing and affective computing. However, the unique brain anatomy of individuals…

Signal Processing · Electrical Eng. & Systems 2024-05-31 Yihang Dong , Xuhang Chen , Yanyan Shen , Michael Kwok-Po Ng , Tao Qian , Shuqiang Wang

Graph neural networks (GNNs) are gaining popularity for processing graph-structured data. In real-world scenarios, graph data within the same dataset can vary significantly in scale. This variability leads to depth-sensitivity, where the…

Machine Learning · Computer Science 2024-11-06 Zelin Yao , Chuang Liu , Xianke Meng , Yibing Zhan , Jia Wu , Shirui Pan , Wenbin Hu
‹ Prev 1 4 5 6 7 8 10 Next ›