中文
相关论文

相关论文: Cross-Modal Manifold Learning for Cross-modal Retr…

200 篇论文

We introduce a novel co-learning paradigm for manifolds naturally equipped with a group action, motivated by recent developments on learning a manifold from attached fibre bundle structures. We utilize a representation theoretic mechanism…

机器学习 · 计算机科学 2019-12-10 Yifeng Fan , Tingran Gao , Zhizhen Zhao

This paper puts forth a novel bi-linear modeling framework for data recovery via manifold-learning and sparse-approximation arguments and considers its application to dynamic magnetic-resonance imaging (dMRI). Each temporal-domain MR image…

图像与视频处理 · 电气工程与系统科学 2020-02-28 Gaurav N. Shetty , Konstantinos Slavakis , Abhishek Bose , Ukash Nakarmi , Gesualdo Scutari , Leslie Ying

Multimodal Contrastive Learning (MCL) advances in aligning different modalities and generating multimodal representations in a joint space. By leveraging contrastive learning across diverse modalities, large-scale multimodal data enhances…

机器学习 · 计算机科学 2025-09-23 Xiaohao Liu , Xiaobo Xia , See-Kiong Ng , Tat-Seng Chua

Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Jiuxiang Gu , Jianfei Cai , Shafiq Joty , Li Niu , Gang Wang

Recent progress has rapidly advanced our understanding of the mechanisms underlying in-context learning in modern attention-based neural networks. However, existing results focus exclusively on unimodal data; in contrast, the theoretical…

机器学习 · 统计学 2026-05-19 Nicholas Barnfield , Subhabrata Sen , Pragya Sur

Learning medical visual representations directly from paired images and reports through multimodal self-supervised learning has emerged as a novel and efficient approach to digital diagnosis in recent years. However, existing models suffer…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Libin Lan , Hongxing Li , Zunhui Xia , Juan Zhou , Xiaofei Zhu , Yongmei Li , Yudong Zhang , Xin Luo

There has been much recent interest in adapting undersampled trajectories in MRI based on training data. In this work, we propose a novel patient-adaptive MRI sampling algorithm based on grouping scans within a training set. Scan-adaptive…

图像与视频处理 · 电气工程与系统科学 2026-02-26 Siddhant Gautam , Angqi Li , Saiprasad Ravishankar

Due to the ever-growing diversity of the data source, multi-modality feature learning has attracted more and more attention. However, most of these methods are designed by jointly learning feature representation from multi-modalities that…

计算机视觉与模式识别 · 计算机科学 2020-06-09 Danfeng Hong , Jocelyn Chanussot , Naoto Yokoya , Jian Kang , Xiao Xiang Zhu

Cross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Yabing Wang , Fan Wang , Jianfeng Dong , Hao Luo

Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human annotations. These joint embeddings enable zero-shot…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Swetha Sirnam , Mamshad Nayeem Rizve , Nina Shvetsova , Hilde Kuehne , Mubarak Shah

The application of compressed sensing (CS)-enabled data reconstruction for accelerating magnetic resonance imaging (MRI) remains a challenging problem. This is due to the fact that the information lost in k-space from the acceleration mask…

图像与视频处理 · 电气工程与系统科学 2023-06-22 Guoyao Shen , Boran Hao , Mengyu Li , Chad W. Farris , Ioannis Ch. Paschalidis , Stephan W. Anderson , Xin Zhang

Cross-modal retrieval across image and text modalities is a challenging task due to its inherent ambiguity: An image often exhibits various situations, and a caption can be coupled with diverse images. Set-based embedding has been studied…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Dongwon Kim , Namyup Kim , Suha Kwak

Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features from text while representing each image with a single feature,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Wei Wu , Xiaomeng Fan , Yuwei Wu , Zhi Gao , Pengxiang Li , Yunde Jia , Mehrtash Harandi

The recent advancements in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively. Building upon this capability, we propose to enable…

多媒体 · 计算机科学 2024-02-19 Yongqi Li , Wenjie Wang , Leigang Qu , Liqiang Nie , Wenjie Li , Tat-Seng Chua

Healthcare applications are inherently multimodal, benefiting greatly from the integration of diverse data sources. However, the modalities available in clinical settings can vary across different locations and patients. A key area that…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Mohammed Amer , Mohamed A. Suliman , Tu Bui , Nuria Garcia , Serban Georgescu

Co-clustering targets on grouping the samples (e.g., documents, users) and the features (e.g., words, ratings) simultaneously. It employs the dual relation and the bilateral information between the samples and features. In many realworld…

机器学习 · 计算机科学 2016-11-18 Ping Li , Jiajun Bu , Chun Chen , Zhanying He , Deng Cai

The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medical foundation models for specific segmentation tasks remains a computational bottleneck. Existing…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Jiaqi Tang , Shaoyang Zhang , Xiaoqi Wang , Jiaying Zhou , Yang Liu , Qingchao Chen

The skeleton of a digital image is a compact representation of its topology, geometry, and scale. It has utility in many computer vision applications, such as image description, segmentation, and registration. However, skeletonization has…

Cross-modal hashing is an important approach for multimodal data management and application. Existing unsupervised cross-modal hashing algorithms mainly rely on data features in pre-trained models to mine their similarity relationships.…

信息检索 · 计算机科学 2022-07-12 Liang Li , Baihua Zheng , Weiwei Sun

Vision-language retrieval is an important multi-modal learning topic, where the goal is to retrieve the most relevant visual candidate for a given text query. Recently, pre-trained models, e.g., CLIP, show great potential on retrieval…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Haojun Jiang , Jianke Zhang , Rui Huang , Chunjiang Ge , Zanlin Ni , Shiji Song , Gao Huang