English
Related papers

Related papers: RE-TRIANGLE: Does TRIANGLE Enable Multimodal Align…

200 papers

This paper presents a deep nonlinear metric learning framework for data visualization on an image dataset. We propose the Triangular Similarity and prove its equivalence to the Cosine Similarity in measuring a data pair. Based on this novel…

Computer Vision and Pattern Recognition · Computer Science 2016-12-28 Lilei Zheng , Ying Zhang , Stefan Duffner , Khalid Idrissi , Christophe Garcia , Atilla Baskurt

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jiale Liu , Haoming Zhou , Yishu Liu , Bingzhi Chen , Yuncheng Jiang

This paper presents a new scalable algorithm for cross-modal similarity preserving retrieval in a learnt manifold space. Unlike existing approaches that compromise between preserving global and local geometries, the proposed technique…

Computer Vision and Pattern Recognition · Computer Science 2016-12-20 Sailesh Conjeti , Anees Kazi , Nassir Navab , Amin Katouzian

The cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different…

Multimedia · Computer Science 2015-12-18 Cuicui Kang , Shengcai Liao , Yonghao He , Jian Wang , Wenjia Niu , Shiming Xiang , Chunhong Pan

Translation averaging aims to recover camera locations from pairwise relative translation directions and is a fundamental component of global Structure-from-Motion pipelines. The problem is challenging because direction measurements contain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Zhekai Fan , Wanze Li , Jinxin Wang , Yunpeng Shi

With the advantage of low storage cost and high efficiency, hashing learning has received much attention in the domain of Big Data. In this paper, we propose a novel unsupervised hashing learning method to cope with this open problem to…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Jun Yu , Xiao-Jun Wu

Current state-of-the-art approaches to cross-modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that attend over all words and objects in an image. While…

Computer Vision and Pattern Recognition · Computer Science 2022-02-22 Gregor Geigle , Jonas Pfeiffer , Nils Reimers , Ivan Vulić , Iryna Gurevych

Contrastive learning has become a standard approach for unsupervised learning from paired data, as demonstrated by CLIP for image-text matching. However, many domains involve more than two modalities and require objectives that capture…

Machine Learning · Computer Science 2026-05-08 Tillmann Rheude , Stefan Hegselmann , Roland Eils , Benjamin Wild

The modern image search system requires semantic understanding of image, and a key yet under-addressed problem is to learn a good metric for measuring the similarity between images. While deep metric learning has yielded impressive…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Jian Wang , Feng Zhou , Shilei Wen , Xiao Liu , Yuanqing Lin

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Runze He , Yiji Cheng , Tiankai Hang , Zhimin Li , Yu Xu , Zijin Yin , Shiyi Zhang , Wenxun Dai , Penghui Du , Ao Ma , Chunyu Wang , Qinglin Lu , Jizhong Han , Jiao Dai

Semantic retrieval, which retrieves semantically matched items given a textual query, has been an essential component to enhance system effectiveness in e-commerce search. In this paper, we study the multimodal retrieval problem, where the…

Information Retrieval · Computer Science 2025-06-26 Zhigong Zhou , Ning Ding , Xiaochuan Fan , Yue Shang , Yiming Qiu , Jingwei Zhuo , Zhiwei Ge , Songlin Wang , Lin Liu , Sulong Xu , Han Zhang

Monocular Metric Depth Estimation (MMDE) is essential for physically intelligent systems, yet accurate depth estimation for underrepresented classes in complex scenes remains a persistent challenge. To address this, we propose RAD, a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Michael Baltaxe , Dan Levi , Sagie Benaim

Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less consideration of cross-lingual alignment during training. Although standardized contrastive…

Computation and Language · Computer Science 2026-04-15 Seungyoon Lee , Minhyuk Kim , Seongtae Hong , Youngjoon Jang , Dongsuk Oh , Heuiseok Lim

Image-text retrieval requires the system to bridge the heterogenous gap between vision and language for accurate retrieval while keeping the network lightweight-enough for efficient retrieval. Existing trade-off solutions mainly study from…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Jiamin Zhuang , Jing Yu , Yang Ding , Xiangyan Qu , Yue Hu

Semantic image synthesis is a challenging task with many practical applications. Albeit remarkable progress has been made in semantic image synthesis with spatially-adaptive normalization and existing methods normalize the feature…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Yupeng Shi , Xiao Liu , Yuxiang Wei , Zhongqin Wu , Wangmeng Zuo

Stereo matching serves as a cornerstone in 3D vision, aiming to establish pixel-wise correspondences between stereo image pairs for depth recovery. Despite remarkable progress driven by deep neural architectures, current models often…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Xianda Guo , Chenming Zhang , Youmin Zhang , Ruilin Wang , Dujun Nie , Wenzhao Zheng , Matteo Poggi , Hao Zhao , Mang Ye , Qin Zou , Long Chen

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversity and improving…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Jiaze Wang , Yi Wang , Ziyu Guo , Renrui Zhang , Donghao Zhou , Guangyong Chen , Anfeng Liu , Pheng-Ann Heng

In this work, we tackle the essential problem of scale inconsistency for self-supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes…

Computer Vision and Pattern Recognition · Computer Science 2021-09-06 Wang Zhao , Shaohui Liu , Yezhi Shu , Yong-Jin Liu

Reranking is a critical component in many information retrieval pipelines. Despite remarkable progress in text-only settings, multimodal reranking remains challenging, particularly when the candidate set contains hybrid text and image…

Information Retrieval · Computer Science 2026-05-26 Yupei Yang , Lin Yang , Wanxi Deng , Lin Qu , Shikui Tu , Lei Xu