English
Related papers

Related papers: INQUIRE: A Natural World Text-to-Image Retrieval B…

200 papers

Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-textual fact-finding. However, evaluating these visual and textual search abilities is still…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yu Zeng , Wenxuan Huang , Zhen Fang , Shuang Chen , Yufan Shen , Yishuo Cai , Xiaoman Wang , Zhenfei Yin , Lin Chen , Zehui Chen , Shiting Huang , Yiming Zhao , Xu Tang , Yao Hu , Philip Torr , Wanli Ouyang , Shaosheng Cao

The ideal form of Visual Question Answering requires understanding, grounding and reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most existing VQA benchmarks are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Kang Chen , Xiangqian Wu

E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a…

Information Retrieval · Computer Science 2026-05-19 Xinyu Sun , Huangyu Dai , Lingtao Mao , Zexin Zheng , Zihan Liang , Ben Chen , Chenyi Lei , Wenwu Ou

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Chenghao Xiao , Isaac Chung , Imene Kerboua , Jamie Stirling , Xin Zhang , Márton Kardos , Roman Solomatin , Noura Al Moubayed , Kenneth Enevoldsen , Niklas Muennighoff

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Pengfei Zhou , Xiaopeng Peng , Jiajun Song , Chuanhao Li , Zhaopan Xu , Yue Yang , Ziyao Guo , Hao Zhang , Yuqi Lin , Yefei He , Lirui Zhao , Shuo Liu , Tianhua Li , Yuxuan Xie , Xiaojun Chang , Yu Qiao , Wenqi Shao , Kaipeng Zhang

Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Wenbo Hu , Jia-Chen Gu , Zi-Yi Dou , Mohsen Fayyaz , Pan Lu , Kai-Wei Chang , Nanyun Peng

While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains largely unknown due to the lack of controlled yet realistic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Changda Zhou , Ziyue Gao , Xueqing Wang , Tingquan Gao , Cheng Cui , Jing Tang , Yi Liu

With the rapid progress of Multimodal LLMs, evaluating their mathematical reasoning capabilities has become an increasingly important research direction. In particular, visual-textual mathematical reasoning serves as a key indicator of an…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Hao Liang , Linzhuang Sun , Minxuan Zhou , Zirong Chen , Meiyi Qiang , Mingan Lin , Tianpeng Li , Fan Yang , Zenan Zhou , Wentao Zhang

This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2024. This challenge is to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Xiaohong Liu , Xiongkuo Min , Guangtao Zhai , Chunyi Li , Tengchuan Kou , Wei Sun , Haoning Wu , Yixuan Gao , Yuqin Cao , Zicheng Zhang , Xiele Wu , Radu Timofte , Fei Peng , Huiyuan Fu , Anlong Ming , Chuanming Wang , Huadong Ma , Shuai He , Zifei Dou , Shu Chen , Huacong Zhang , Haiyi Xie , Chengwei Wang , Baoying Chen , Jishen Zeng , Jianquan Yang , Weigang Wang , Xi Fang , Xiaoxin Lv , Jun Yan , Tianwu Zhi , Yabin Zhang , Yaohui Li , Yang Li , Jingwen Xu , Jianzhao Liu , Yiting Liao , Junlin Li , Zihao Yu , Yiting Lu , Xin Li , Hossein Motamednia , S. Farhad Hosseini-Benvidi , Fengbin Guan , Ahmad Mahmoudi-Aznaveh , Azadeh Mansouri , Ganzorig Gankhuyag , Kihwan Yoon , Yifang Xu , Haotian Fan , Fangyuan Kong , Shiling Zhao , Weifeng Dong , Haibing Yin , Li Zhu , Zhiling Wang , Bingchen Huang , Avinab Saha , Sandeep Mishra , Shashank Gupta , Rajesh Sureddi , Oindrila Saha , Luigi Celona , Simone Bianco , Paolo Napoletano , Raimondo Schettini , Junfeng Yang , Jing Fu , Wei Zhang , Wenzhi Cao , Limei Liu , Han Peng , Weijun Yuan , Zhan Li , Yihang Cheng , Yifan Deng , Haohui Li , Bowen Qu , Yao Li , Shuqing Luo , Shunzhou Wang , Wei Gao , Zihao Lu , Marcos V. Conde , Xinrui Wang , Zhibo Chen , Ruling Liao , Yan Ye , Qiulin Wang , Bing Li , Zhaokun Zhou , Miao Geng , Rui Chen , Xin Tao , Xiaoyu Liang , Shangkun Sun , Xingyuan Ma , Jiaze Li , Mengduo Yang , Haoran Xu , Jie Zhou , Shiding Zhu , Bohan Yu , Pengfei Chen , Xinrui Xu , Jiabin Shen , Zhichao Duan , Erfan Asadi , Jiahe Liu , Qi Yan , Youran Qu , Xiaohui Zeng , Lele Wang , Renjie Liao

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, which severely…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Jia Wang , Jie Hu , Xiaoqi Ma , Hanghang Ma , Yanbing Zeng , Xiaoming Wei

We investigate knowledge retrieval with multi-modal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called…

Computation and Language · Computer Science 2023-06-02 Man Luo , Zhiyuan Fang , Tejas Gokhale , Yezhou Yang , Chitta Baral

This work introduces ILIAS, a new test dataset for Instance-Level Image retrieval At Scale. It is designed to evaluate the ability of current and future foundation models and retrieval techniques to recognize particular objects. The key…

The rapid development of large language models (LLMs) and large vision models (LVMs) have propelled the evolution of multi-modal AI systems, which have demonstrated the remarkable potential for industrial applications by emulating…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Di Jin , Xing Liu , Yu Liu , Jia Qing Yap , Andrea Wong , Adriana Crespo , Qi Lin , Zhiyuan Yin , Qiang Yan , Ryan Ye

An outstanding image-text retrieval model depends on high-quality labeled data. While the builders of existing image-text retrieval datasets strive to ensure that the caption matches the linked image, they cannot prevent a caption from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Xu Yan , Chunhui Ai , Ziqiang Cao , Min Cao , Sujian Li , Wenjie Li , Guohong Fu

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate information from multiple…

Information Retrieval · Computer Science 2022-09-29 Cheng-An Hsieh , Cheng-Ping Hsieh , Pu-Jen Cheng

Semi-iNat is a challenging dataset for semi-supervised classification with a long-tailed distribution of classes, fine-grained categories, and domain shifts between labeled and unlabeled data. This dataset is behind the second iteration of…

Computer Vision and Pattern Recognition · Computer Science 2021-06-24 Jong-Chyi Su , Subhransu Maji

We present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 140k unique images annotated with ground truth by human raters…

Computer Vision and Pattern Recognition · Computer Science 2018-08-02 Troy Chinen , Johannes Ballé , Chunhui Gu , Sung Jin Hwang , Sergey Ioffe , Nick Johnston , Thomas Leung , David Minnen , Sean O'Malley , Charles Rosenberg , George Toderici

Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on feature enhancement…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Jie Wang , Joemon M. Jose

Multimodal encoders have pushed the boundaries of visual document retrieval, matching textual query tokens directly to image patches and achieving state-of-the-art performance on public benchmarks. Recent models relying on this paradigm…

Computation and Language · Computer Science 2026-04-08 Omri Uzan , Asaf Yehudai , Roi pony , Eyal Shnarch , Ariel Gera

The goal of text-to-video retrieval is to search large databases for relevant videos based on text queries. Existing methods have progressed to handling explicit queries where the visual content of interest is described explicitly; however,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yiqing Shen , Chenxiao Fan , Chenjia Li , Mathias Unberath