English
Related papers

Related papers: MoRE: Multi-Modal Contrastive Pre-training with Tr…

200 papers

Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-target pairs across different modalities. Yet, despite its…

Information Retrieval · Computer Science 2026-04-03 Geonmo Gu , Byeongho Heo , Jaemyung Yu , Jaehui Hwang , Taekyung Kim , Sangmin Lee , HeeJae Jun , Yoohoon Kang , Sangdoo Yun , Dongyoon Han

Electrocardiogram (ECG) diagnosis remains challenging due to limited labeled data and the need to capture subtle yet clinically meaningful variations in rhythm and morphology. We present CREMA (Contrastive Regularized Masked Autoencoder), a…

Machine Learning · Computer Science 2025-08-22 Junho Song , Jong-Hwan Jang , DongGyun Hong , Joon-myoung Kwon , Yong-Yeon Jo

Multi-modal fusion approaches aim to integrate information from different data sources. Unlike natural datasets, such as in audio-visual applications, where samples consist of "paired" modalities, data in healthcare is often collected…

Image and Video Processing · Electrical Eng. & Systems 2023-03-03 Nasir Hayat , Krzysztof J. Geras , Farah E. Shamout

Recent advancements in non-invasive detection of cardiac hemodynamic instability (CHDI) primarily focus on applying machine learning techniques to a single data modality, e.g. cardiac magnetic resonance imaging (MRI). Despite their…

Electrocardiography (ECG) analysis is crucial for cardiac diagnosis, yet existing foundation models often fail to capture the periodicity and diverse features required for varied clinical tasks. We propose ECG-MoE, a hybrid architecture…

Artificial Intelligence · Computer Science 2026-03-06 Yuhao Xu , Xiaoda Wang , Yi Wu , Wei Jin , Xiao Hu , Carl Yang

Multi-modal entity alignment aims to identify equivalent entities between two different multi-modal knowledge graphs, which consist of structural triples and images associated with entities. Most previous works focus on how to utilize and…

Computation and Language · Computer Science 2022-09-05 Zhenxi Lin , Ziheng Zhang , Meng Wang , Yinghui Shi , Xian Wu , Yefeng Zheng

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Yiming Sun , Bing Cao , Pengfei Zhu , Qinghua Hu

Both functional and structural magnetic resonance imaging (fMRI and sMRI) are widely used for the diagnosis of mental disorder. However, combining complementary information from these two modalities is challenging due to their…

Image and Video Processing · Electrical Eng. & Systems 2024-04-02 Ziyu Zhou , Anton Orlichenko , Gang Qu , Zening Fu , Vince D Calhoun , Zhengming Ding , Yu-Ping Wang

A vital and rapidly growing application, remote sensing offers vast yet sparsely labeled, spatially aligned multimodal data; this makes self-supervised learning algorithms invaluable. We present CROMA: a framework that combines contrastive…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Anthony Fuller , Koreen Millard , James R. Green

Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often directly fuse the two modalities for cross-reference during the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Yeong-Joon Ju , Ho-Joong Kim , Seong-Whan Lee

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Frank Li , Bardia Khosravi , Mohammadreza Chavoshi , Young Seok Jeon , Theo Dapamede , Hari Trivedi , Janice Newsome , Judy Gichoya

Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. While some recent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Stefanos Koutoupis , Michaela Areti Zervou , Konstantinos Kontras , Maarten De Vos , Panagiotis Tsakalides , Grigorios Tsagkatakis

Multi-modal magnetic resonance (MR) imaging provides great potential for diagnosing and analyzing brain gliomas. In clinical scenarios, common MR sequences such as T1, T2 and FLAIR can be obtained simultaneously in a single scanning…

Image and Video Processing · Electrical Eng. & Systems 2022-03-10 Ziqi Huang , Li Lin , Pujin Cheng , Linkai Peng , Xiaoying Tang

With the rapid advancement of multi-modal large language models (MLLMs) in recent years, the foundational Contrastive Language-Image Pretraining (CLIP) framework has been successfully extended to MLLMs, enabling more powerful and universal…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Youze Xue , Dian Li , Gang Liu

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Representations (CALM), an…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-09 Vin Sachidananda , Shao-Yen Tseng , Erik Marchi , Sachin Kajarekar , Panayiotis Georgiou

Data is one of the essential ingredients to power deep learning research. Small datasets, especially specific to medical institutes, bring challenges to deep learning training stage. This work aims to develop a practical deep multimodal…

Machine Learning · Computer Science 2019-02-26 Faik Aydin , Maggie Zhang , Michelle Ananda-Rajah , Gholamreza Haffari

Automated generation of clinically accurate radiology reports can improve patient care. Previous report generation methods that rely on image captioning models often generate incoherent and incorrect text due to their lack of relevant…

Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation.We propose a novel cross-modal framework that uses MedCLIP as both…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Qianhao Han , Junyi Liu , Zengchang Qin , Zheng Zheng

Automatic detection of multimodal fake news has gained a widespread attention recently. Many existing approaches seek to fuse unimodal features to produce multimodal news representations. However, the potential of powerful cross-modal…

Machine Learning · Computer Science 2023-08-14 Longzheng Wang , Chuang Zhang , Hongbo Xu , Yongxiu Xu , Xiaohan Xu , Siqi Wang

Multi-modal affect recognition models leverage complementary information in different modalities to outperform their uni-modal counterparts. However, due to the unavailability of modality-specific sensors or data, multi-modal models may not…

Image and Video Processing · Electrical Eng. & Systems 2021-08-03 Vandana Rajan , Alessio Brutti , Andrea Cavallaro
‹ Prev 1 8 9 10 Next ›