English
Related papers

Related papers: Cross-Modal Contrastive Representation Learning fo…

200 papers

Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pose data, which eliminates background perturbation but…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Kepeng Wu , Zecheng Li , Hezhen Hu , Wengang Zhou , Houqiang Li

Learning medical visual representations directly from paired images and reports through multimodal self-supervised learning has emerged as a novel and efficient approach to digital diagnosis in recent years. However, existing models suffer…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Libin Lan , Hongxing Li , Zunhui Xia , Juan Zhou , Xiaofei Zhu , Yongmei Li , Yudong Zhang , Xin Luo

Text-to-image generation and image captioning are recently emerged as a new experimental paradigm to assess machine intelligence. They predict continuous quantity accompanied by their sampling techniques in the generation, making evaluation…

Computer Vision and Pattern Recognition · Computer Science 2022-05-27 Jin-Hwa Kim , Yunji Kim , Jiyoung Lee , Kang Min Yoo , Sang-Woo Lee

Multi-modal magnetic resonance imaging (MRI) is essential for providing complementary information about brain anatomy and pathology, leading to more accurate diagnoses. However, obtaining high-quality multi-modal MRI in a clinical setting…

Image and Video Processing · Electrical Eng. & Systems 2025-04-15 Minjoo Lim , Bogyeong Kang , Tae-Eui Kam

Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Yabing Wang , Le Wang , Qiang Zhou , Zhibin Wang , Hao Li , Gang Hua , Wei Tang

In deepfake detection, the varying degrees of compression employed by social media platforms pose significant challenges for model generalization and reliability. Although existing methods have progressed from single-modal to multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ching-Yi Lai , Chih-Yu Jian , Pei-Cheng Chuang , Chia-Ming Lee , Chih-Chung Hsu , Chiou-Ting Hsu , Chia-Wen Lin

Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-level semantic information captured during large-scale…

Computer Vision and Pattern Recognition · Computer Science 2020-07-28 Lei Shi , Kai Shuang , Shijie Geng , Peng Su , Zhengkai Jiang , Peng Gao , Zuohui Fu , Gerard de Melo , Sen Su

Multi-modal generation has been widely explored in recent years. Current research directions involve generating text based on an image or vice versa. In this paper, we propose a new task called CIGLI: Conditional Image Generation from…

Computer Vision and Pattern Recognition · Computer Science 2021-08-23 Xiaopeng Lu , Lynnette Ng , Jared Fernandez , Hao Zhu

Multi-modal keyphrase generation aims to produce a set of keyphrases that represent the core points of the input text-image pair. In this regard, dominant methods mainly focus on multi-modal fusion for keyphrase generation. Nevertheless,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yifan Dong , Suhang Wu , Fandong Meng , Jie Zhou , Xiaoli Wang , Jianxin Lin , Jinsong Su

Heterogeneous gap among different modalities emerges as one of the critical issues in modern AI problems. Unlike traditional uni-modal cases, where raw features are extracted and directly measured, the heterogeneous nature of cross modal…

Information Retrieval · Computer Science 2015-11-19 Aiwen Jiang , Hanxi Li , Yi Li , Mingwen Wang

Could we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio…

Sound · Computer Science 2022-02-23 Ziyu Wang , Dejing Xu , Gus Xia , Ying Shan

Image-text multimodal representation learning aligns data across modalities and enables important medical applications, e.g., image classification, visual grounding, and cross-modal retrieval. In this work, we establish a connection between…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Peiqi Wang , William M. Wells , Seth Berkowitz , Steven Horng , Polina Golland

Text-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture, necessitating the integration of retrieval methods. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mengdan Zhu , Senhao Cheng , Guangji Bai , Yifei Zhang , Liang Zhao

Self-Supervised Contrastive Learning has proven effective in deriving high-quality representations from unlabeled data. However, a major challenge that hinders both unimodal and multimodal contrastive learning is feature suppression, a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jihai Zhang , Xiang Lan , Xiaoye Qu , Yu Cheng , Mengling Feng , Bryan Hooi

Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose…

Multi-modal magnetic resonance imaging (MRI) provides rich, complementary information for analyzing diseases. However, the practical challenges of acquiring multiple MRI modalities, such as cost, scan time, and safety considerations, often…

Image and Video Processing · Electrical Eng. & Systems 2024-09-16 Zhaohu Xing , Sicheng Yang , Sixiang Chen , Tian Ye , Yijun Yang , Jing Qin , Lei Zhu

Contrastive learning has been shown to produce generalizable representations of audio and visual data by maximizing the lower bound on the mutual information (MI) between different views of an instance. However, obtaining a tight lower…

Machine Learning · Computer Science 2021-04-20 Shuang Ma , Zhaoyang Zeng , Daniel McDuff , Yale Song

Age estimation of face images is a crucial task with various practical applications in areas such as video surveillance and Internet access control. While deep learning-based age estimation frameworks, e.g., convolutional neural network…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Yuntao Shou , Xiangyong Cao , Deyu Meng

Multimodal sentiment analysis has become an increasingly popular research area as the demand for multimodal online content is growing. For multimodal sentiment analysis, words can have different meanings depending on the linguistic context…

Computation and Language · Computer Science 2022-09-16 Junghun Kim , Jihie Kim

Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training, as introducing more modalities can greatly complicate model…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Ludan Ruan , Anwen Hu , Yuqing Song , Liang Zhang , Sipeng Zheng , Qin Jin
‹ Prev 1 4 5 6 7 8 10 Next ›