English
Related papers

Related papers: Modular Multimodal Classification Without Fine-Tun…

200 papers

Cancer prognosis is a critical task that involves predicting patient outcomes and survival rates. To enhance prediction accuracy, previous studies have integrated diverse data modalities, such as clinical notes, medical images, and genomic…

Machine Learning · Computer Science 2025-02-04 Jie Peng , Shuang Zhou , Longwei Yang , Yiran Song , Mohan Zhang , Kaixiong Zhou , Feng Xie , Mingquan Lin , Rui Zhang , Tianlong Chen

Multiple instance learning (MIL) significantly reduced annotation costs via bag-level weak labels for large-scale images, such as histopathological whole slide images (WSIs). However, its adaptability to continual tasks with minimal…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Byung Hyun Lee , Wongi Jeong , Woojae Han , Kyoungbun Lee , Se Young Chun

The prevalence of large-scale multimodal datasets presents unique challenges in assessing dataset quality. We propose a two-step method to analyze multimodal datasets, which leverages a small seed of human annotation to map each multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-07-11 Netta Madvil , Yonatan Bitton , Roy Schwartz

Classification of social media data is an important approach in understanding user behavior on the Web. Although information on social media can be of different modalities such as texts, images, audio or videos, traditional approaches in…

Computation and Language · Computer Science 2017-08-08 Chi Thang Duong , Remi Lebret , Karl Aberer

Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other…

Machine Learning · Computer Science 2024-11-05 Zirun Guo , Tao Jin , Jingyuan Chen , Zhou Zhao

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model…

Sound · Computer Science 2025-06-24 Angelos-Nikolaos Kanatas , Charilaos Papaioannou , Alexandros Potamianos

Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities are inconsistent.…

Artificial Intelligence · Computer Science 2025-10-14 Zhiyuan Han , Beier Zhu , Yanlong Xu , Peipei Song , Xun Yang

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-modal reasoning remains underexplored, with conflicting reports on whether added modalities help or…

Computation and Language · Computer Science 2026-05-01 Yucheng Wang , Yifan Hou , Aydin Javadov , Mubashara Akhtar , Mrinmaya Sachan

Visual foundation models (VFMs) have become increasingly popular due to their state-of-the-art performance. However, interpretability remains crucial for critical applications. In this sense, self-explainable models (SEM) aim to provide…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Hugues Turbé , Mina Bjelogrlic , Gianmarco Mengaldo , Christian Lovis

In this paper, we explore a less-studied yet practically important problem: how to efficiently and effectively adapt multiple ($>$2) multimodal foundation models (MFMs) for the sequential recommendation task. To this end, we propose a…

Information Retrieval · Computer Science 2025-09-16 Junchen Fu , Yongxin Ni , Joemon M. Jose , Ioannis Arapakis , Kaiwen Zheng , Youhua Li , Xuri Ge

Code comment classification is a critical task for automated software documentation and analysis. In the context of the NLBSE'26 Tool Competition, we present LoRA-MME, a Multi-Model Ensemble architecture utilizing Parameter-Efficient…

Software Engineering · Computer Science 2026-04-16 Md Akib Haider , Ahsan Bulbul , Nafis Fuad Shahid , Aimaan Ahmed , Mohammad Ishrak Abedin

Multimodal semantic segmentation enhances model robustness by exploiting cross-modal complementarities. However, existing methods often suffer from imbalanced modal dependencies, where overall performance degrades significantly once a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Jiaqi Tan , Xu Zheng , Fangyu Li , Yang Liu

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder…

Artificial Intelligence · Computer Science 2023-12-15 Haoyu Zhang , Yu Wang , Guanghao Yin , Kejun Liu , Yuanyuan Liu , Tianshu Yu

Multi-modal pre-trained models efficiently extract and fuse features from different modalities with low memory requirements for fine-tuning. Despite this efficiency, their application in disease diagnosis is under-explored. A significant…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Zhiyi Shi , Junsik Kim , Wanhua Li , Yicong Li , Hanspeter Pfister

In clinical practice, crossmodal information including medical images and tabular data is essential for disease diagnosis. There exists a significant modality gap between these data types, which obstructs advancements in crossmodal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Tianling Liu , Hongying Liu , Fanhua Shang , Lequan Yu , Tong Han , Liang Wan

As medical diagnoses increasingly leverage multimodal data, machine learning models are expected to effectively fuse heterogeneous information while remaining robust to missing modalities. In this work, we propose a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Yi Gu , Kuniaki Saito , Jiaxin Ma

The missing modality problem poses a fundamental challenge in multimodal sentiment analysis, significantly degrading model accuracy and generalization in real world scenarios. Existing approaches primarily improve robustness through prompt…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Rongfei Chen , Tingting Zhang , Xiaoyu Shen , Wei Zhang

Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial expression recognition (FER) methods struggle to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Fuyan Ma , Yiran He , Bin Sun , Shutao Li

Robust multimodal systems must remain effective when some modalities are noisy, degraded, or unreliable. Existing multimodal fusion methods often learn modality selection jointly with representation learning, making it difficult to…

Artificial Intelligence · Computer Science 2026-03-31 Roland Bertin-Johannet , Lara Scipio , Leopold Maytié , Rufin VanRullen

The development of large language models (LLMs) has expanded to multi-modal systems capable of processing text, images, and speech within a unified framework. Training these models demands significantly larger datasets and computational…

Computation and Language · Computer Science 2025-05-09 Weixin Liang , Lili Yu , Liang Luo , Srinivasan Iyer , Ning Dong , Chunting Zhou , Gargi Ghosh , Mike Lewis , Wen-tau Yih , Luke Zettlemoyer , Xi Victoria Lin