English
Related papers

Related papers: TransModality: An End2End Fusion Method with Trans…

200 papers

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Xuechen Wang , Shiwan Zhao , Haoqin Sun , Hui Wang , Jiaming Zhou , Yong Qin

In this paper we propose a deep learning method for performing attributed-based music-to-image translation. The proposed method is applied for synthesizing visual stories according to the sentiment expressed by songs. The generated images…

Computer Vision and Pattern Recognition · Computer Science 2019-12-13 Nikolaos Passalis , Stavros Doropoulos

Multimodal sentiment analysis is a fundamental problem in the field of affective computing. Although significant progress has been made in cross-modal interaction, it remains a challenge due to the insufficient reference context in…

Multimedia · Computer Science 2025-08-12 Xianbing Zhao , Shengzun Yang , Buzhou Tang , Ronghuan Jiang

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over…

Artificial Intelligence · Computer Science 2024-08-21 Chunting Zhou , Lili Yu , Arun Babu , Kushal Tirumala , Michihiro Yasunaga , Leonid Shamis , Jacob Kahn , Xuezhe Ma , Luke Zettlemoyer , Omer Levy

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text…

Computation and Language · Computer Science 2018-02-26 Yue Gu , Shuhong Chen , Ivan Marsic

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Zexu Pan , Zhaojie Luo , Jichen Yang , Haizhou Li

Multi-modal affective computing aims to automatically recognize and interpret human attitudes from diverse data sources such as images and text, thereby enhancing human-computer interaction and emotion understanding. Existing approaches…

Computation and Language · Computer Science 2025-06-10 Yuanhe Tian , Pengsen Cheng , Guoqing Jin , Lei Zhang , Yan Song

Multimodal sentiment analysis enhances conventional sentiment analysis, which traditionally relies solely on text, by incorporating information from different modalities such as images, text, and audio. This paper proposes a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Taoxu Zhao , Meisi Li , Kehao Chen , Liye Wang , Xucheng Zhou , Kunal Chaturvedi , Mukesh Prasad , Ali Anaissi , Ali Braytee

Multimodal sentiment analysis is drawing an increasing amount of attention these days. It enables mining of opinions in video reviews which are now available aplenty on online platforms. However, multimodal sentiment analysis has only a few…

Computation and Language · Computer Science 2017-04-14 Haohan Wang , Aaksha Meghawat , Louis-Philippe Morency , Eric P. Xing

Multimodal Sentiment Analysis (MSA) is critical for human-computer interaction but faces challenges when the modalities are incomplete or missing. Existing methods often assume pre-defined missing modalities or fixed missing rates, limiting…

Human-Computer Interaction · Computer Science 2025-11-24 Liling Li , Guoyang Xu , Xiongri Shen , Zhifei Xu , Yanbo Zhang , Zhiguo Zhang , Zhenxi Song

Emotion Prediction in Conversation (EPC) aims to forecast the emotions of forthcoming utterances by utilizing preceding dialogues. Previous EPC approaches relied on simple context modeling for emotion extraction, overlooking fine-grained…

Multimedia · Computer Science 2024-08-09 Haoxiang Shi , Ziqi Liang , Jun Yu

Multimodal speech emotion recognition (SER) has emerged as pivotal for improving human-machine interaction. Researchers are increasingly leveraging both speech and textual information obtained through automatic speech recognition (ASR) to…

Human-Computer Interaction · Computer Science 2025-09-24 Jiajun He , Xiaohan Shi , Cheng-Hung Hu , Jinyi Mi , Xingfeng Li , Tomoki Toda

Multimodal Sentiment Analysis (MSA) utilizes multimodal data to infer the users' sentiment. Previous methods focus on equally treating the contribution of each modality or statically using text as the dominant modality to conduct…

Computation and Language · Computer Science 2024-10-08 Xinyu Feng , Yuming Lin , Lihua He , You Li , Liang Chang , Ya Zhou

Multimodal sentiment analysis has been studied under the assumption that all modalities are available. However, such a strong assumption does not always hold in practice, and most of multimodal fusion models may fail when partial modalities…

Machine Learning · Computer Science 2022-05-02 Jiandian Zeng , Tianyi Liu , Jiantao Zhou

Integration of multimodal information from various sources has been shown to boost the performance of machine learning models and thus has received increased attention in recent years. Often such models use deep modality-specific networks…

Machine Learning · Computer Science 2022-11-22 Shiv Shankar , Laure Thompson , Madalina Fiterau

The fusion technique is the key to the multimodal emotion recognition task. Recently, cross-modal attention-based fusion methods have demonstrated high performance and strong robustness. However, cross-modal attention suffers from redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Feng Liu , Ziwang Fu , Yunlong Wang , Qijian Zheng

Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker…

Sound · Computer Science 2025-11-19 Xiao Li , Kotaro Funakoshi , Manabu Okumura

Speech emotion recognition is a challenging task and an important step towards more natural human-machine interaction. We show that pre-trained language models can be fine-tuned for text emotion recognition, achieving an accuracy of 69.5%…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Verena Heusser , Niklas Freymuth , Stefan Constantin , Alex Waibel

This paper aims to demonstrate the importance and feasibility of fusing multimodal information for emotion recognition. It introduces a multimodal framework for emotion understanding by fusing the information from visual facial features and…

Artificial Intelligence · Computer Science 2023-06-06 Puneet Kumar , Xiaobai Li

Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms. However, these models are facing challenges of overfitting with limited labels and low model generalization abilities. In this…

Sound · Computer Science 2021-09-02 Hang Li , Yu Kang , Tianqiao Liu , Wenbiao Ding , Zitao Liu