English
Related papers

Related papers: Self-MI: Efficient Multimodal Fusion via Self-Supe…

200 papers

We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jaeyoo Park , Bohyung Han

Human Activity Recognition is a field of research where input data can take many forms. Each of the possible input modalities describes human behaviour in a different way, and each has its own strengths and weaknesses. We explore the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Razvan Brinzea , Bulat Khaertdinov , Stylianos Asteriadis

Multi-modality image fusion is a technique that combines information from different sensors or modalities, enabling the fused image to retain complementary features from each modality, such as functional highlights and texture details.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Zixiang Zhao , Haowen Bai , Jiangshe Zhang , Yulun Zhang , Kai Zhang , Shuang Xu , Dongdong Chen , Radu Timofte , Luc Van Gool

Due to the severe lack of labeled data, existing methods of medical visual question answering usually rely on transfer learning to obtain effective image feature representation and use cross-modal fusion of visual and linguistic features to…

Multimedia · Computer Science 2021-05-04 Haifan Gong , Guanqi Chen , Sishuo Liu , Yizhou Yu , Guanbin Li

RGB-Infrared (RGB-IR) multimodal perception is fundamental to embodied multimedia systems operating in complex physical environments. Although recent cross-modal fusion methods have advanced RGB-IR detection, the optimization dynamics…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Xianhui Liu , Siqi Jiang , Yi Xie , Yuqing Lin , Siao Liu

Multimodal sentiment analysis is an important research area that predicts speaker's sentiment tendency through features extracted from textual, visual and acoustic modalities. The central challenge is the fusion method of the multimodal…

Computation and Language · Computer Science 2020-09-29 Zilong Wang , Zhaohong Wan , Xiaojun Wan

Fusing data from multiple modalities provides more information to train machine learning systems. However, it is prohibitively expensive and time-consuming to label each modality with a large amount of data, which leads to a crucial problem…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Xinwei Sun , Yilun Xu , Peng Cao , Yuqing Kong , Lingjing Hu , Shanghang Zhang , Yizhou Wang

Learning to collaborate is critical in Multi-Agent Reinforcement Learning (MARL). Previous works promote collaboration by maximizing the correlation of agents' behaviors, which is typically characterized by Mutual Information (MI) in…

Multiagent Systems · Computer Science 2023-02-23 Pengyi Li , Hongyao Tang , Tianpei Yang , Xiaotian Hao , Tong Sang , Yan Zheng , Jianye Hao , Matthew E. Taylor , Wenyuan Tao , Zhen Wang , Fazl Barez

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance changes caused by…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Abhinav Valada , Rohit Mohan , Wolfram Burgard

Multimodal representation learning is a challenging task in which previous work mostly focus on either uni-modality pre-training or cross-modality fusion. In fact, we regard modeling multimodal representation as building a skyscraper, where…

Computation and Language · Computer Science 2024-08-15 Ronghao Lin , Haifeng Hu

Multimodal entity linking (MEL) task, which aims at resolving ambiguous mentions to a multimodal knowledge graph, has attracted wide attention in recent years. Though large efforts have been made to explore the complementary effect among…

Artificial Intelligence · Computer Science 2023-07-20 Pengfei Luo , Tong Xu , Shiwei Wu , Chen Zhu , Linli Xu , Enhong Chen

The convergence of cross-modal adversarial learning and physics-driven methods represents a cutting-edge direction for tackling challenges in complex multi-modal tasks and scientific computing. This review focuses on systematically…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hana Satou , Alan Mitkiy

In this paper, we introduce InSQuAD, designed to enhance the performance of In-Context Learning (ICL) models through Submodular Mutual Information} (SMI) enforcing Quality and Diversity among in-context exemplars. InSQuAD achieves this…

Machine Learning · Computer Science 2025-08-29 Souradeep Nanda , Anay Majee , Rishabh Iyer

Utilizing multi-modal data enhances scene understanding by providing complementary semantic and geometric information. Existing methods fuse features or distill knowledge from multiple modalities into a unified representation, improving…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jialei Chen , Xu Zheng , Danda Pani Paudel , Luc Van Gool , Hiroshi Murase , Daisuke Deguchi

Multimodal image-tabular learning is gaining attention, yet it faces challenges due to limited labeled data. While earlier work has applied self-supervised learning (SSL) to unlabeled data, its task-agnostic nature often results in learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Siyi Du , Xinzhe Luo , Declan P. O'Regan , Chen Qin

Multimodal human action understanding is a significant problem in computer vision, with the central challenge being the effective utilization of the complementarity among diverse modalities while maintaining model efficiency. However, most…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Hongsong Wang , Heng Fei , Bingxuan Dai , Jie Gui

High annotation costs are a substantial bottleneck in applying modern deep learning architectures to clinically relevant medical use cases, substantiating the need for novel algorithms to learn from unlabeled data. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Aiham Taleb , Matthias Kirchler , Remo Monti , Christoph Lippert

Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Ameera Bawazir , Kebin Wu , Wenbin Li

Unsupervised visible infrared person re-identification (USVI-ReID) is a challenging retrieval task that aims to retrieve cross-modality pedestrian images without using any label information. In this task, the large cross-modality variance…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Zhizhong Zhang , Jiangming Wang , Xin Tan , Yanyun Qu , Junping Wang , Yong Xie , Yuan Xie

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary unpaired…

Machine Learning · Computer Science 2025-10-10 Sharut Gupta , Shobhita Sundaram , Chenyu Wang , Stefanie Jegelka , Phillip Isola