English
Related papers

Related papers: CoVR-2: Automatic Data Construction for Composed V…

200 papers

Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility in unseen multimodal retrieval. Recent LLM-based advances focus on generating the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Miaoge Li , Dongsheng Wang , Zening Sun , Jinsen Zhang , Wenhan Luo , Jingcai Guo

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a compositional query, consisting of a reference image and a modifying text-without relying on annotated training data. Existing approaches often generate a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Rong-Cheng Tu , Wenhao Sun , Hanzhe You , Yingjie Wang , Jiaxing Huang , Li Shen , Dacheng Tao

Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Zixu Li , Yupeng Hu , Zhiwei Chen , Shiqi Zhang , Qinlei Huang , Zhiheng Fu , Yinwei Wei

Learning the similarity between remote sensing (RS) images forms the foundation for content-based RS image retrieval (CBIR). Recently, deep metric learning approaches that map the semantic similarity of images into an embedding (metric)…

Computer Vision and Pattern Recognition · Computer Science 2021-11-10 Gencer Sumbul , Mahdyar Ravanbakhsh , Begüm Demir

The relations expressed in user queries are vital for cross-modal information retrieval. Relation-focused cross-modal retrieval aims to retrieve information that corresponds to these relations, enabling effective retrieval across different…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Yan Gong , Georgina Cosma , Axel Finke

Generating dense multiview images from text prompts is crucial for creating high-fidelity 3D assets. Nevertheless, existing methods struggle with space-view correspondences, resulting in sparse and low-quality outputs. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Bonan Li , Zicheng Zhang , Xingyi Yang , Xinchao Wang

Vision Language Models (VLMs) have achieved remarkable breakthroughs in the field of remote sensing in recent years. Synthetic Aperture Radar (SAR) imagery, with its all-weather capability, is essential in remote sensing, yet the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Yiguo He , Xinjun Cheng , Junjie Zhu , Chunping Qiu , Jun Wang , Xichuan Zhang , Qiangjuan Huang , Ke Yang

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Yiming Ju , Jijin Hu , Zhengxiong Luo , Haoge Deng , hanyu Zhao , Li Du , Chengwei Wu , Donglin Hao , Xinlong Wang , Tengfei Pan

Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Tomohisa Takeda , Yu-Chieh Lin , Yuji Nozawa , Youyang Ng , Osamu Torii , Yusuke Matsui

Content-based image retrieval (CBIR) is one of the most active research areas in multimedia information retrieval. Given a query image, the task is to search relevant images in a repository. Low level features like color, texture, and shape…

Computer Vision and Pattern Recognition · Computer Science 2018-12-12 Asheet Kumar , Shivam Choudhary , Vaibhav Singh Khokhar , Vikas Meena , Chiranjoy Chattopadhyay

In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images. Since most…

Computation and Language · Computer Science 2017-05-03 Yuya Yoshikawa , Yutaro Shigeto , Akikazu Takeuchi

While there exists a lot of work on explainable complaint mining, articulating user concerns through text or video remains a significant challenge, often leaving issues unresolved. Users frequently struggle to express their complaints…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Sarmistha Das , R E Zera Marveen Lyngkhoi , Kirtan Jain , Vinayak Goyal , Sriparna Saha , Manish Gupta

Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual with modification text to retrieve the desired target. Despite this shared structure,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Haokun Wen , Xuemeng Song , Haoyu Zhang , Xiangyu Zhao , Weili Guan , Liqiang Nie

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding, which hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Houlun Chen , Xin Wang , Hong Chen , Zeyang Zhang , Wei Feng , Bin Huang , Jia Jia , Wenwu Zhu

Text-to-image retrieval aims to find the relevant images based on a text query, which is important in various use-cases, such as digital libraries, e-commerce, and multimedia databases. Although Multimodal Large Language Models (MLLMs)…

Information Retrieval · Computer Science 2024-04-04 Zijun Long , Xuri Ge , Richard Mccreadie , Joemon Jose

The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge. Human-generated queries for video datasets `in the wild' vary a lot in terms of degree of specificity,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-17 Yang Liu , Samuel Albanie , Arsha Nagrani , Andrew Zisserman

Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically require manual YouTube…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Timing Yang , Sucheng Ren , Alan Yuille , Feng Wang

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Xin Gu , Ming Li , Libo Zhang , Fan Chen , Longyin Wen , Tiejian Luo , Sijie Zhu

Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visually realistic images.…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Soyeon Caren Han , Siqu Long , Siwen Luo , Kunze Wang , Josiah Poon

Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Chunhui Zhang , Yiren Jian , Zhongyu Ouyang , Soroush Vosoughi