English
Related papers

Related papers: A Closer Look at Multimodal Representation Collaps…

200 papers

Learning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, and the discriminative…

Machine Learning · Computer Science 2022-12-06 Sijie Mai , Ying Zeng , Haifeng Hu

Multi-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the problem of the neglect of inter-feature relationships, high-frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xiaoli Zhang , Liying Wang , Libo Zhao , Xiongfei Li , Siwei Ma

The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Yi Wang , Conrad M Albrecht , Nassim Ait Ali Braham , Chenying Liu , Zhitong Xiong , Xiao Xiang Zhu

This work focuses on learning useful and robust deep world models using multiple, possibly unreliable, sensors. We find that current methods do not sufficiently encourage a shared representation between modalities; this can cause poor…

Machine Learning · Computer Science 2021-07-07 Kaiqi Chen , Yong Lee , Harold Soh

Effectively integrating diverse sensory modalities is crucial for robotic manipulation. However, the typical approach of feature concatenation is often suboptimal: dominant modalities such as vision can overwhelm sparse but critical signals…

Noise has always been nonnegligible trouble in object detection by creating confusion in model reasoning, thereby reducing the informativeness of the data. It can lead to inaccurate recognition due to the shift in the observed pattern, that…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Xinyu Zhang , Zhiwei Li , Zhenhong Zou , Xin Gao , Yijin Xiong , Dafeng Jin , Jun Li , Huaping Liu

We formalize and study a phenomenon called feature collapse that makes precise the intuitive idea that entities playing a similar role in a learning task receive similar representations. As feature collapse requires a notion of task, we…

Machine Learning · Computer Science 2023-05-26 Thomas Laurent , James H. von Brecht , Xavier Bresson

Multimodal deep learning, especially vision-language models, have gained significant traction in recent years, greatly improving performance on many downstream tasks, including content moderation and violence detection. However, standard…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Zhuokai Zhao , Harish Palani , Tianyi Liu , Lena Evans , Ruth Toner

Multimodal learning systems often encounter challenges related to modality imbalance, where a dominant modality may overshadow others, thereby hindering the learning of weak modalities. Conventional approaches often force weak modalities to…

Machine Learning · Computer Science 2025-10-27 Baoquan Gong , Xiyuan Gao , Pengfei Zhu , Qinghua Hu , Bing Cao

Correlations between factors of variation are prevalent in real-world data. Exploiting such correlations may increase predictive performance on noisy data; however, often correlations are not robust (e.g., they may change between domains,…

Machine Learning · Computer Science 2022-12-26 Christina M. Funke , Paul Vicol , Kuan-Chieh Wang , Matthias Kümmerer , Richard Zemel , Matthias Bethge

Purpose High dimensional, multimodal data can nowadays be analyzed by huge deep neural networks with little effort. Several fusion methods for bringing together different modalities have been developed. Given the prevalence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Christian Gapp , Elias Tappeiner , Martin Welk , Karl Fritscher , Elke Ruth Gizewski , Rainer Schubert

Multi-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are…

Machine Learning · Computer Science 2021-03-10 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Zhuohang Dang , Minnan Luo , Jihong Wang , Chengyou Jia , Haochen Han , Herun Wan , Guang Dai , Xiaojun Chang , Jingdong Wang

Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space…

Machine Learning · Computer Science 2023-11-07 Vinitra Swamy , Malika Satayeva , Jibril Frej , Thierry Bossy , Thijs Vogels , Martin Jaggi , Tanja Käser , Mary-Anne Hartley

Using multiple spatial modalities has been proven helpful in improving semantic segmentation performance. However, there are several real-world challenges that have yet to be addressed: (a) improving label efficiency and (b) enhancing…

Computer Vision and Pattern Recognition · Computer Science 2023-04-24 Harsh Maheshwari , Yen-Cheng Liu , Zsolt Kira

Multi-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain, highlighting…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Guanzhou Ke , Bo Wang , Xiaoli Wang , Shengfeng He

Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downstream clinical tasks. However, most existing MFMs suffer from information ambiguity that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yihang Liu , Longzhen Yang , Jiaxiong Yang , Ying Wen , Lianghua He , Heng Tao Shen

Large models have demonstrated exceptional generalization capabilities in computer vision and natural language processing. Recent efforts have focused on enhancing these models with multimodal processing abilities. However, addressing the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Hao Sun , Yu Song

Multimodal sentiment analysis (MSA) draws increasing attention with the availability of multimodal data. The boost in performance of MSA models is mainly hindered by two problems. On the one hand, recent MSA works mostly focus on learning…

Machine Learning · Computer Science 2021-11-17 Ying Zeng , Sijie Mai , Haifeng Hu

Multi-modal learning has achieved remarkable success by integrating information from various modalities, achieving superior performance in tasks like recognition and retrieval compared to uni-modal approaches. However, real-world scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xiaohao Liu , Xiaobo Xia , Zhuo Huang , See-Kiong Ng , Tat-Seng Chua