中文
相关论文

相关论文: Cross-Modality Gated Attention Fusion for Multimod…

200 篇论文

Multimodal sensors provide complementary information to develop accurate machine-learning methods for human activity recognition (HAR), but introduce significantly higher computational load, which reduces efficiency. This paper proposes an…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Ziqi Gao , Yuntao Wang , Jianguo Chen , Junliang Xing , Shwetak Patel , Xin Liu , Yuanchun Shi

Multi-modal learning has shown exceptional performance in various tasks, especially in medical applications, where it integrates diverse medical information for comprehensive diagnostic evidence. However, there still are several challenges…

机器学习 · 计算机科学 2024-11-19 Lin Fan , Yafei Ou , Cenyang Zheng , Pengyu Dai , Tamotsu Kamishima , Masayuki Ikebe , Kenji Suzuki , Xun Gong

We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM uses self-attention to learn the latent concepts in scene…

计算机视觉与模式识别 · 计算机科学 2018-09-24 Kyung-Min Kim , Seong-Ho Choi , Jin-Hwa Kim , Byoung-Tak Zhang

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important…

音频与语音处理 · 电气工程与系统科学 2023-02-07 Jiachen Luo , Huy Phan , Joshua Reiss

We examine the impact of concept-informed supervision on multimodal video interpretation models using MOByGaze, a dataset containing human-annotated explanatory concepts. We introduce Concept Modality Specific Datasets (CMSDs), which…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Elisa Ancarani , Julie Tores , Lucile Sassatelli , Rémy Sun , Hui-Yin Wu , Frédéric Precioso

In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to extract and combine…

计算与语言 · 计算机科学 2021-09-17 Wei Han , Hui Chen , Soujanya Poria

In this paper, we introduce an Adaptive Graph Signal Processing with Dynamic Semantic Alignment (AGSP DSA) framework to perform robust multimodal data fusion over heterogeneous sources, including text, audio, and images. The requested…

计算机视觉与模式识别 · 计算机科学 2026-01-27 KV Karthikeya , Ashok Kumar Das , Shantanu Pal , Vivekananda Bhat K , Arun Sekar Rajasekaran

Modern Web systems such as social media and e-commerce contain rich contents expressed in images and text. Leveraging information from multi-modalities can improve the performance of machine learning tasks such as classification and…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Huidong Liu , Shaoyuan Xu , Jinmiao Fu , Yang Liu , Ning Xie , Chien-Chih Wang , Bryan Wang , Yi Sun

Query-based moment localization is a new task that localizes the best matched segment in an untrimmed video according to a given sentence query. In this localization task, one should pay more attention to thoroughly mine visual and…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Daizong Liu , Xiaoye Qu , Xiao-Yang Liu , Jianfeng Dong , Pan Zhou , Zichuan Xu

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

With the rapid growth of AI-generated content (AIGC) across domains such as music, video, and literature, the demand for emotionally aware recommendation systems has become increasingly important. Traditional recommender systems primarily…

信息检索 · 计算机科学 2025-12-15 Zheqi Hu , Xuanjing Chen , Jinlin Hu

Incorporating additional sensory modalities such as tactile and audio into foundational robotic models poses significant challenges due to the curse of dimensionality. This work addresses this issue through modality selection. We propose a…

机器人学 · 计算机科学 2025-04-22 Jiawei Jiang , Kei Ota , Devesh K. Jha , Asako Kanezaki

Multimodal emotion recognition in conversation (MER) aims to accurately identify emotions in conversational utterances by integrating multimodal information. Previous methods usually treat multimodal information as equal quality and employ…

多媒体 · 计算机科学 2024-11-18 Xiaofei Zhu , Jiawei Cheng , Zhou Yang , Zhuo Chen , Qingyang Wang , Jianfeng Yao

We introduce AdaptiSent, a new framework for Multimodal Aspect-Based Sentiment Analysis (MABSA) that uses adaptive cross-modal attention mechanisms to improve sentiment classification and aspect term extraction from both text and images.…

计算与语言 · 计算机科学 2025-07-18 S M Rafiuddin , Sadia Kamal , Mohammed Rakib , Arunkumar Bagavathi , Atriya Sen

Video sentiment analysis as a decision-making process is inherently complex, involving the fusion of decisions from multiple modalities and the so-caused cognitive biases. Inspired by recent advances in quantum cognition, we show that the…

计算与语言 · 计算机科学 2021-05-20 Dimitris Gkoumas , Qiuchi Li , Shahram Dehdashti , Massimo Melucci , Yijun Yu , Dawei Song

Learning modality-fused representations and processing unaligned multimodal sequences are meaningful and challenging in multimodal emotion recognition. Existing approaches use directional pairwise attention or a message hub to fuse…

计算机视觉与模式识别 · 计算机科学 2021-12-06 Ziwang Fu , Feng Liu , Hanyang Wang , Siyuan Shen , Jiahao Zhang , Jiayin Qi , Xiangling Fu , Aimin Zhou

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applications, such as…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Qingyun Fang , Zhaokui Wang

Aggregating multi-modality data to obtain reliable data representation attracts more and more attention. Recent studies demonstrate that Transformer models usually work well for multi-modality tasks. Existing Transformers generally either…

计算机视觉与模式识别 · 计算机科学 2023-03-17 Xixi Wang , Xiao Wang , Bo Jiang , Jin Tang , Bin Luo

Sentiment analysis and emotion recognition in videos are challenging tasks, given the diversity and complexity of the information conveyed in different modalities. Developing a highly competent framework that effectively addresses the…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Prasad Chaudhari , Aman Kumar , Chandravardhan Singh Raghaw , Mohammad Zia Ur Rehman , Nagendra Kumar

As an important multimodal sentiment analysis task, Joint Multimodal Aspect-Sentiment Analysis (JMASA), aiming to jointly extract aspect terms and their associated sentiment polarities from the given text-image pairs, has gained increasing…

计算与语言 · 计算机科学 2024-05-24 Yaxin Liu , Yan Zhou , Ziming Li , Jinchuan Zhang , Yu Shang , Chenyang Zhang , Songlin Hu