English
Related papers

Related papers: CoMIR: Contrastive Multimodal Image Representation…

200 papers

Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities.…

Machine Learning · Computer Science 2023-10-20 Zehan Wang , Yang Zhao , Xize Cheng , Haifeng Huang , Jiageng Liu , Li Tang , Linjun Li , Yongqi Wang , Aoxiong Yin , Ziang Zhang , Zhou Zhao

This work proposes a multimodal diffeomorphic registration method using Neural Ordinary Differential Equations (Neural ODEs). Nonrigid registration algorithms exhibit tradeoffs between their accuracy, the computational complexity of their…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Salvador Rodriguez-Sanz , Monica Hernandez

Magnetic Resonance Imaging (MRI) typically recruits multiple sequences (defined here as "modalities"). As each modality is designed to offer different anatomical and functional clinical information, there are evident disparities in the…

Image and Video Processing · Electrical Eng. & Systems 2022-03-09 Chengjia Wang , Guang Yang , Giorgos Papanastasiou

Image registration techniques usually assume that the images to be registered are of a certain type (e.g. single- vs. multi-modal, 2D vs. 3D, rigid vs. deformable) and there lacks a general method that can work for data under all…

Image and Video Processing · Electrical Eng. & Systems 2025-01-28 Quang Luong Nhat Nguyen , Ruiming Cao , Laura Waller

Cross-modal medical image-report retrieval task plays a significant role in clinical diagnosis and various medical generative tasks. Eliminating heterogeneity between different modalities to enhance semantic consistency is the key challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Zeqiang Wei , Kai Jin , Xiuzhuang Zhou

Code search, framed as information retrieval (IR), underpins modern software engineering and increasingly powers retrieval-augmented generation (RAG), improving code discovery, reuse, and the reliability of LLM-based coding. Yet existing…

Software Engineering · Computer Science 2026-04-20 Jiahui Geng , Qing Li , Fengyu Cai , Fakhri Karray

Deformable shape representations, parameterized by deformations relative to a given template, have proven effective for improved image analysis tasks. However, their broader applicability is hindered by two major challenges. First, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Tonmoy Hossain , Miaomiao Zhang

Human perception integrates multiple modalities, such as vision, hearing, and language, into a unified understanding of the surrounding reality. While recent multimodal models have achieved significant progress by aligning pairs of…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Giordano Cicchetti , Eleonora Grassucci , Luigi Sigillo , Danilo Comminiello

Few-shot classification requires deep neural networks to learn generalized representations only from limited training images, which is challenging but significant in low-data regimes. Recently, CLIP-based methods have shown promising…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Renrui Zhang , Bohao Li , Wei Zhang , Hao Dong , Hongsheng Li , Peng Gao , Yu Qiao

Traditional model-based image reconstruction (MBIR) methods combine forward and noise models with simple object priors. Recent application of deep learning methods for image reconstruction provides a successful data-driven approach to…

Image and Video Processing · Electrical Eng. & Systems 2023-11-22 Ling Chen , Zhishen Huang , Yong Long , Saiprasad Ravishankar

Effective representation of Regions of Interest (ROI) and independent alignment of these ROIs can significantly enhance the performance of deformable medical image registration (DMIR). However, current learning-based DMIR methods have…

Image and Video Processing · Electrical Eng. & Systems 2025-06-25 Xinke Ma , Yongsheng Pan , Qingjie Zeng , Mengkang Lu , Bolysbek Murat Yerzhanuly , Bazargul Matkerim , Yong Xia

We propose a coercive approach to simultaneously register and segment multi-modal images which share similar spatial structure. Registration is done at the region level to facilitate data fusion while avoiding the need for interpolation.…

Computer Vision and Pattern Recognition · Computer Science 2015-11-19 Yu-Hui Chen , Dennis Wei , Gregory Newstadt , Jeffrey Simmons , Alfred Hero

Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. While some recent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Stefanos Koutoupis , Michaela Areti Zervou , Konstantinos Kontras , Maarten De Vos , Panagiotis Tsakalides , Grigorios Tsagkatakis

The success of many computer vision tasks lies in the ability to exploit the interdependency between different image modalities such as intensity and depth. Fusing corresponding information can be achieved on several levels, and one…

Computer Vision and Pattern Recognition · Computer Science 2014-06-26 Martin Kiechle , Tim Habigt , Simon Hawe , Martin Kleinsteuber

Recent works in medical image registration have proposed the use of Implicit Neural Representations, demonstrating performance that rivals state-of-the-art learning-based methods. However, these implicit representations need to be optimized…

Image and Video Processing · Electrical Eng. & Systems 2023-10-04 Louis D. van Harten , Jaap Stoker , Ivana Išgum

Brain imaging classification is commonly approached from two perspectives: modeling the full image volume to capture global anatomical context, or constructing ROI-based graphs to encode localized and topological interactions. Although both…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Wei Liang , Lifang He

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two…

Artificial Intelligence · Computer Science 2026-02-25 Ziwei Niu , Hao Sun , Shujun Bian , Xihong Yang , Lanfen Lin , Yuxin Liu , Yueming Jin

There are a wide range of applications that involve multi-modal data, such as cross-modal retrieval, visual question-answering, and image captioning. Such applications are primarily dependent on aligned distributions of the different…

Machine Learning · Computer Science 2020-05-26 Vishaal Udandarao , Abhishek Maiti , Deepak Srivatsav , Suryatej Reddy Vyalla , Yifang Yin , Rajiv Ratn Shah

The massive availability of cameras results in a wide variability of imaging conditions, producing large intra-class variations and a significant performance drop if heterogeneous images are compared for person recognition. However, as…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Fernando Alonso-Fernandez , Kiran B. Raja , R. Raghavendra , Cristoph Busch , Josef Bigun , Ruben Vera-Rodriguez , Julian Fierrez

The image-text retrieval task aims to retrieve relevant information from a given image or text. The main challenge is to unify multimodal representation and distinguish fine-grained differences across modalities, thereby finding similar…

Multimedia · Computer Science 2024-05-20 Ziyu Gong , Chengcheng Mai , Yihua Huang
‹ Prev 1 3 4 5 6 7 10 Next ›