English
Related papers

Related papers: Learning TFIDF Enhanced Joint Embedding for Recipe…

200 papers

Joint embedding (JE) is a way to encode multi-modal data into a vector space where text remains as the grounding key and other modalities like image are to be anchored with such keys. Meme is typically an image with embedded text onto it.…

Machine Learning · Computer Science 2021-12-06 Nethra Gunti , Sathyanarayanan Ramamoorthy , Parth Patwa , Amitava Das

Multimodal medical image fusion (MMIF) aims to integrate images from different modalities to produce a comprehensive image that enhances medical diagnosis by accurately depicting organ structures, tissue textures, and metabolic information.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Tao Luo , Weihua Xu

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

Sound · Computer Science 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Over the past two decades, recommendation systems (RSs) have used machine learning (ML) solutions to recommend items, e.g., movies, books, and restaurants, to clients of a business or an online platform. Recipe recommendation, however, has…

Information Retrieval · Computer Science 2023-08-10 Ali Pesaranghader , Touqir Sajed

Most existing learning-based multi-modality image fusion (MMIF) methods suffer from significant structure inconsistency due to their inappropriate usage of structural features at the semantic level. To alleviate these issues, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Qiao Yang , Yu Zhang , Yutong Chen , Jian Zhang , Shunli Zhang

Low-dimensional embeddings for data from disparate sources play critical roles in multi-modal machine learning, multimedia information retrieval, and bioinformatics. In this paper, we propose a supervised dimensionality reduction method…

Machine Learning · Computer Science 2021-01-15 Yanjun Li , Bihan Wen , Hao Cheng , Yoram Bresler

Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approaches tackle this problem using implicit feature interaction…

Computer Vision and Pattern Recognition · Computer Science 2020-10-02 Shaofei Huang , Tianrui Hui , Si Liu , Guanbin Li , Yunchao Wei , Jizhong Han , Luoqi Liu , Bo Li

In this paper, we present a cross-modal recipe retrieval framework, Transformer-based Network for Large Batch Training (TNLBT), which is inspired by ACME~(Adversarial Cross-Modal Embedding) and H-T~(Hierarchical Transformer). TNLBT aims to…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Jing Yang , Junwen Chen , Keiji Yanai

Direct computer vision based-nutrient content estimation is a demanding task, due to deformation and occlusions of ingredients, as well as high intra-class and low inter-class variability between meal classes. In order to tackle these…

Information Retrieval · Computer Science 2019-11-06 Matthias Fontanellaz , Stergios Christodoulidis , Stavroula Mougiakakou

Conventional multimodal data integration methods provide a comprehensive assessment of the shared or unique structure within each individual data type but suffer from several limitations such as the inability to handle high-dimensional data…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Matthew Drexler , Benjamin Risk , James J Lah , Suprateek Kundu , Deqiang Qiu

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Yale Song , Mohammad Soleymani

Understanding expressed sentiment and emotions are two crucial factors in human multimodal language. This paper describes a Transformer-based joint-encoding (TBJE) for the task of Emotion Recognition and Sentiment Analysis. In addition to…

Computation and Language · Computer Science 2020-08-11 Jean-Benoit Delbrouck , Noé Tits , Mathilde Brousmiche , Stéphane Dupont

Multi-modal image fusion (MMIF) enhances the information content of the fused image by combining the unique as well as common features obtained from different modality sensor images, improving visualization, object detection, and many more…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Gargi Panda , Soumitra Kundu , Saumik Bhattacharya , Aurobinda Routray

Multi-modal entity alignment (MMEA) is essential for enhancing knowledge graphs and improving information retrieval and question-answering systems. Existing methods often focus on integrating modalities through their complementarity but…

Artificial Intelligence · Computer Science 2024-10-21 Wei Ai , Wen Deng , Hongyi Chen , Jiayi Du , Tao Meng , Yuntao Shou

Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yao-Hung Hubert Tsai , Liang-Kang Huang , Ruslan Salakhutdinov

Remote sensing image interpretation plays a critical role in environmental monitoring, urban planning, and disaster assessment. However, acquiring high-quality labeled data is often costly and time-consuming. To address this challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Wang , Guanzhou Chen , Xiaodong Zhang , Chenxi Liu , Jiaqi Wang , Xiaoliang Tan , Wenchao Guo , Qingyuan Yang , Kaiqi Zhang

Building interpretation from remote sensing imagery primarily involves two fundamental tasks: building extraction and change detection. However, most existing methods address these tasks independently, overlooking their inherent correlation…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Dehua Huo , Weida Zhan , Jinxin Guo , Depeng Zhu , Yu Chen , YiChun Jiang , Yueyi Han , Deng Han , Jin Li

Multi-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xilai Li , Xiaosong Li , Tao Ye , Xiaoqi Cheng , Wuyang Liu , Haishu Tan

Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using…

Multimedia · Computer Science 2024-12-10 Shanti Stewart , Kleanthis Avramidis , Tiantian Feng , Shrikanth Narayanan

Sequential recommendation (SR) systems excel at capturing users' dynamic preferences by leveraging their interaction histories. Most existing SR systems assign a single embedding vector to each item to represent its features, adopting…

Information Retrieval · Computer Science 2026-01-21 Mingrui Liu , Sixiao Zhang , Cheng Long