English
Related papers

Related papers: Hierachical Delta-Attention Method for Multimodal …

200 papers

Multimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic. Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), multimodal systems…

Machine Learning · Computer Science 2021-09-10 Pierre Colombo , Emile Chapuis , Matthieu Labeau , Chloe Clavel

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment…

Machine Learning · Computer Science 2018-02-06 Minghai Chen , Sen Wang , Paul Pu Liang , Tadas Baltrušaitis , Amir Zadeh , Louis-Philippe Morency

Transformers have advanced the field of natural language processing (NLP) on a variety of important tasks. At the cornerstone of the Transformer architecture is the multi-head attention (MHA) mechanism which models pairwise interactions…

Computation and Language · Computer Science 2021-06-01 Lin Zheng , Zhiyong Wu , Lingpeng Kong

While Diffusion Models excel in text-to-image synthesis, they often suffer from concept omission when synthesizing complex multi-instance scenes. Existing training-free methods attempt to resolve this by rescaling attention maps, which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Zitong Wang , Zijun Shen , Haohao Xu , Zhengjie Luo , Weibin Wu

Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data…

Computation and Language · Computer Science 2019-06-04 Yao-Hung Hubert Tsai , Shaojie Bai , Paul Pu Liang , J. Zico Kolter , Louis-Philippe Morency , Ruslan Salakhutdinov

Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual contents. Existing state-of-the-art approaches use…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Zhou Yu , Yuhao Cui , Jun Yu , Dacheng Tao , Qi Tian

We consider the problem of Visual Question Answering (VQA). Given an image and a free-form, open-ended, question, expressed in natural language, the goal of VQA system is to provide accurate answer to this question with respect to the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-07 Tanzila Rahman , Shih-Han Chou , Leonid Sigal , Giuseppe Carenini

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional…

Robotics · Computer Science 2024-09-19 Youssef Mohamed , Severin Lemaignan , Arzu Guneysu , Patric Jensfelt , Christian Smith

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

Computation and Language · Computer Science 2021-01-27 Gaurav Sahu , Olga Vechtomova

This study focuses on how different modalities of human communication can be used to distinguish between healthy controls and subjects with schizophrenia who exhibit strong positive symptoms. We developed a multi-modal schizophrenia…

Signal Processing · Electrical Eng. & Systems 2024-04-22 Gowtham Premananth , Yashish M. Siriwardena , Philip Resnik , Carol Espy-Wilson

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

Sound · Computer Science 2023-09-29 R. Gnana Praveen , Jahangir Alam

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet how they internally integrate visual and textual information remains poorly understood. To bridge this gap, we perform a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Shezheng Song , Shasha Li , Jie Yu

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Zhengyao Lv , Tianlin Pan , Chenyang Si , Zhaoxi Chen , Wangmeng Zuo , Ziwei Liu , Kwan-Yee K. Wong

Micro-expression, for its high objectivity in emotion detection, has emerged to be a promising modality in affective computing. Recently, deep learning methods have been successfully introduced into the micro-expression recognition area.…

Computer Vision and Pattern Recognition · Computer Science 2019-08-28 Chongyang Wang , Min Peng , Tao Bi , Tong Chen

Depression, a prevalent and serious mental health issue, affects approximately 3.8\% of the global population. Despite the existence of effective treatments, over 75\% of individuals in low- and middle-income countries remain untreated,…

Computation and Language · Computer Science 2024-07-19 Shengjie Li , Yinhao Xiao

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In…

Computer Vision and Pattern Recognition · Computer Science 2021-02-10 Linwei Ye , Mrigank Rochan , Zhi Liu , Xiaoqin Zhang , Yang Wang

Effective multimodal fusion requires mechanisms that can capture complex cross-modal dependencies while remaining computationally scalable for real-world deployment. Existing audio-visual fusion approaches face a fundamental trade-off:…

Multimedia · Computer Science 2026-02-03 Mohamed Saleh , Zahra Ahmadi

Hand gestures are a primary output of the human motor system, yet the decoding of their neuromuscular signatures remains a bottleneck for basic neuroscience and assistive technologies such as prosthetics. Traditional human-machine interface…

Machine Learning · Computer Science 2025-06-23 Eion Tyacke , Kunal Gupta , Jay Patel , Rui Li

Due to its ability to accurately predict emotional state using multimodal features, audiovisual emotion recognition has recently gained more interest from researchers. This paper proposes two methods to predict emotional attributes from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-22 Bagus Tris Atmaja , Masato Akagi