English
Related papers

Related papers: Transformer Decoders with MultiModal Regularizatio…

200 papers

With the exponential surge in diverse multi-modal data, traditional uni-modal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged,…

Information Retrieval · Computer Science 2024-10-01 Tianshi Wang , Fengling Li , Lei Zhu , Jingjing Li , Zheng Zhang , Heng Tao Shen

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Zhao Yang , Jiaqi Wang , Yansong Tang , Kai Chen , Hengshuang Zhao , Philip H. S. Torr

Multimodal deep neural networks enhance deep comprehension by integrating diverse data modalities. Data from different modalities are typically projected into a shared latent space for similarity computation, but this process is resource…

Machine Learning · Computer Science 2026-05-19 Alberto Presta , Grzegorz Stefanski , Michal Byra , Krzysztof Arendt

Analog circuit design relies heavily on reusing existing intellectual property (IP), yet searching across heterogeneous representations such as SPICE netlists, schematics, and functional descriptions remains challenging. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yihan Wang , Lei Li , Yao Lai , Jing Wang , Yan Lu

Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval…

Computer Vision and Pattern Recognition · Computer Science 2018-06-14 Jiuxiang Gu , Jianfei Cai , Shafiq Joty , Li Niu , Gang Wang

This paper pays close attention to the cross-modality visible-infrared person re-identification (VI Re-ID) task, which aims to match pedestrian samples between visible and infrared modes. In order to reduce the modality-discrepancy between…

Computer Vision and Pattern Recognition · Computer Science 2022-02-10 Guangwei Gao , Hao Shao , Fei Wu , Meng Yang , Yi Yu

Food recommendation systems serve as pivotal components in the realm of digital lifestyle services, designed to assist users in discovering recipes and food items that resonate with their unique dietary predilections. Typically, multi-modal…

Information Retrieval · Computer Science 2025-02-28 Yixin Zhang , Xin Zhou , Qianwen Meng , Fanglin Zhu , Yonghui Xu , Zhiqi Shen , Lizhen Cui

Significant work has been conducted in the domain of food computing, yet these studies typically focus on single tasks such as t2t (instruction generation from food titles and ingredients), i2t (recipe generation from food images), or t2i…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Peiyu Li , Xiaobao Huang , Yijun Tian , Nitesh V. Chawla

It is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures. We present a Multi-modal Semantics enhanced Joint Embedding…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Zhongwei Xie , Ling Liu , Yanzhao Wu , Lin Li , Luo Zhong

Cross-modal retrieval aims to learn discriminative and modal-invariant features for data from different modalities. Unlike the existing methods which usually learn from the features extracted by offline networks, in this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Longlong Jing , Elahe Vahdani , Jiaxing Tan , Yingli Tian

The cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different…

Multimedia · Computer Science 2015-12-18 Cuicui Kang , Shengcai Liao , Yonghao He , Jian Wang , Wenjia Niu , Shiming Xiang , Chunhong Pan

Cross-modal retrieval is the task of retrieving samples of a given modality by using queries of a different one. Due to the wide range of practical applications, the problem has been mainly focused on the vision and language case, e.g. text…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Jorge Sánchez , Rodrigo Laguna

We deal with the problem of learning the underlying disentangled latent factors that are shared between the paired bi-modal data in cross-modal retrieval. Our assumption is that the data in both modalities are complex, structured, and high…

Machine Learning · Computer Science 2020-12-02 Minyoung Kim , Ricardo Guerrero , Vladimir Pavlovic

Cross-modal 3D retrieval is a critical yet challenging task, aiming to achieve bi-directional retrieval between 3D and text modalities. Current methods predominantly rely on a certain 3D representation (e.g., point cloud), with few…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Junlong Ren , Hao Wang

Cross-modal retrieval is an important functionality in modern search engines, as it increases the user experience by allowing queries and retrieved objects to pertain to different modalities. In this paper, we focus on the image-sentence…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Nicola Messina , Giuseppe Amato , Fabrizio Falchi , Claudio Gennaro , Stéphane Marchand-Maillet

Pre-training on large scale unlabelled datasets has shown impressive performance improvements in the fields of computer vision and natural language processing. Given the advent of large-scale instructional video datasets, a common strategy…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Valentin Gabeur , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Representation learning for sketch-based image retrieval has mostly been tackled by learning embeddings that discard modality-specific information. As instances from different modalities can often provide complementary information…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Abhra Chaudhuri , Massimiliano Mancini , Yanbei Chen , Zeynep Akata , Anjan Dutta

Encoded representations from a pretrained deep learning model (e.g., BERT text embeddings, penultimate CNN layer activations of an image) convey a rich set of features beneficial for information retrieval. Embeddings for a particular…

Machine Learning · Computer Science 2023-04-24 Hyunjin Choi , Hyunjae Lee , Seongho Joe , Youngjune L. Gwon

Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) suffers from heavy computation due to the self-attention…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Fengyuan Shi , Ruopeng Gao , Weilin Huang , Limin Wang

Cross-modal retrieval across image and text modalities is a challenging task due to its inherent ambiguity: An image often exhibits various situations, and a caption can be coupled with diverse images. Set-based embedding has been studied…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Dongwon Kim , Namyup Kim , Suha Kwak