English
Related papers

Related papers: Towards Bridging the Cross-modal Semantic Gap for …

200 papers

In multimodal learning, CLIP has emerged as the de-facto approach for mapping different modalities into a shared latent space by bringing semantically similar representations closer while pushing apart dissimilar ones. However, CLIP-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Eleonora Grassucci , Giordano Cicchetti , Danilo Comminiello

Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful…

Computer Vision and Pattern Recognition · Computer Science 2025-02-07 Marco Mistretta , Alberto Baldrati , Lorenzo Agnolucci , Marco Bertini , Andrew D. Bagdanov

As one of the main solutions to the information overload problem, recommender systems are widely used in daily life. In the recent emerging micro-video recommendation scenario, micro-videos contain rich multimedia information, involving…

Information Retrieval · Computer Science 2022-05-31 Breda Lim , Shubhi Bansal , Ahmed Buru , Kayla Manthey

While recommender systems with multi-modal item representations (image, audio, and text), have been widely explored, learning recommendations from multi-modal user interactions (e.g., clicks and speech) remains an open problem. We study the…

Information Retrieval · Computer Science 2024-05-08 Simone Borg Bruun , Krisztian Balog , Maria Maistro

Multimodal Recommender Systems aim to improve recommendation accuracy by integrating heterogeneous content, such as images and textual metadata. While effective, it remains unclear whether their gains stem from true multimodal understanding…

Information Retrieval · Computer Science 2025-08-07 Claudio Pomo , Matteo Attimonelli , Danilo Danese , Fedelucio Narducci , Tommaso Di Noia

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world application. In this work, we investigate how contrastive…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Binxu Li , Yuhui Zhang , Xiaohan Wang , Weixin Liang , Ludwig Schmidt , Serena Yeung-Levy

Multimodal recommendation has emerged as an effective paradigm for enhancing collaborative filtering by incorporating heterogeneous content modalities. Existing multimodal recommenders predominantly focus on reinforcing cross-modal…

Information Retrieval · Computer Science 2026-03-03 Hao Zhan , Yihui Wang , Yonghui Yang , Danyang Yue , Yu Wang , Pengyang Shao , Fei Shen , Fei Liu , Le Wu

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

Computation and Language · Computer Science 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

Multimedia recommendation aims to fuse the multi-modal information of items for feature enrichment to improve the recommendation performance. However, existing methods typically introduce multi-modal information based on collaborative…

Information Retrieval · Computer Science 2023-07-07 Haokai Ma , Zhuang Qi , Xinxin Dong , Xiangxian Li , Yuze Zheng , Xiangxu Mengand Lei Meng

In recent years, the rapid growth of online multimedia services, such as e-commerce platforms, has necessitated the development of personalised recommendation approaches that can encode diverse content about each item. Indeed, modern…

Information Retrieval · Computer Science 2023-11-06 Zixuan Yi , Zijun Long , Iadh Ounis , Craig Macdonald , Richard Mccreadie

Cross-modal retrieval is the task of retrieving samples of a given modality by using queries of a different one. Due to the wide range of practical applications, the problem has been mainly focused on the vision and language case, e.g. text…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Jorge Sánchez , Rodrigo Laguna

In multimodal learning, CLIP has been recognized as the \textit{de facto} method for learning a shared latent space across multiple modalities, placing similar representations close to each other and moving them away from dissimilar ones.…

Machine Learning · Computer Science 2026-01-27 Eleonora Grassucci , Giordano Cicchetti , Emanuele Frasca , Aurelio Uncini , Danilo Comminiello

Multimodal learning plays a critical role in e-commerce recommendation platforms today, enabling accurate recommendations and product understanding. However, existing vision-language models, such as CLIP, face key challenges in e-commerce…

Information Retrieval · Computer Science 2025-07-24 Ramin Giahi , Kehui Yao , Sriram Kollipara , Kai Zhao , Vahid Mirjalili , Jianpeng Xu , Topojoy Biswas , Evren Korpeoglu , Kannan Achan

Contrastive Language--Image Pre-training (CLIP) has manifested remarkable improvements in zero-shot classification and cross-modal vision-language tasks. Yet, from a geometrical point of view, the CLIP embedding space has been found to have…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Sedigheh Eslami , Gerard de Melo

Contrastive Language-Image Pre-training (CLIP) has drawn increasing attention recently for its transferable visual representation learning. However, due to the semantic gap within datasets, CLIP's pre-trained image-text alignment becomes…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Longtian Qiu , Renrui Zhang , Ziyu Guo , Ziyao Zeng , Zilu Guo , Yafeng Li , Guangnan Zhang

With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner,…

Information Retrieval · Computer Science 2025-02-24 Wenyu Zhang , Jie Luo , Xinming Zhang , Yuan Fang

Multimodal learning has recently gained significant popularity, demonstrating impressive performance across various zero-shot classification tasks and a range of perceptive and generative applications. Models such as Contrastive…

Machine Learning · Computer Science 2026-02-16 Can Yaras , Siyi Chen , Peng Wang , Qing Qu

Many recommendation systems limit user inputs to text strings or behavior signals such as clicks and purchases, and system outputs to a list of products sorted by relevance. With the advent of generative AI, users have come to expect richer…

Multimodal recommendation focuses primarily on effectively exploiting both behavioral and multimodal information for the recommendation task. However, most existing models suffer from the following issues when fusing information from two…

Information Retrieval · Computer Science 2024-09-10 Kangning Zhang , Yingjie Qin , Jiarui Jin , Yifan Liu , Ruilong Su , Weinan Zhang , Yong Yu

Multimodal recommender systems work by augmenting the representation of the products in the catalogue through multimodal features extracted from images, textual descriptions, or audio tracks characterising such products. Nevertheless, in…

Information Retrieval · Computer Science 2024-04-01 Daniele Malitesta , Emanuele Rossi , Claudio Pomo , Fragkiskos D. Malliaros , Tommaso Di Noia
‹ Prev 1 2 3 10 Next ›