English
Related papers

Related papers: INQUIRE: A Natural World Text-to-Image Retrieval B…

200 papers

The spreading of AI-generated images (AIGI), driven by advances in generative AI, poses a significant threat to information security and public trust. Existing AIGI detectors, while effective against images in clean laboratory settings,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Cheng Xia , Manxi Lin , Jiexiang Tan , Xiaoxiong Du , Yang Qiu , Junjun Zheng , Xiangheng Kong , Yuning Jiang , Bo Zheng

Recently, the Metaverse is becoming increasingly attractive, with millions of users accessing the many available virtual worlds. However, how do users find the one Metaverse which best fits their current interests? So far, the search…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Ali Abdari , Alex Falcon , Giuseppe Serra

With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant research challenge. Current methods, including those based on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Xinli Yue , JianHui Sun , Junda Lu , Liangchao Yao , Fan Xia , Tianyi Wang , Fengyun Rao , Jing Lyu , Yuetang Deng

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluations rely largely on structured-text tables or clean…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Junzhe Huang , Xiaoxiao Sun , Yan Yang , Yuxuan Hou , Ruotian Zhang , Sirui Li , Hehe Fan , Serena Yeung-Levy , Xin Yu

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

Visual search is an essential part of almost any everyday human goal-directed interaction with the environment. Nowadays, several algorithms are able to predict gaze positions during simple observation, but few models attempt to simulate…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 F. Travi , G. Ruarte , G. Bujia , J. E. Kamienkowski

Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated…

This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Zheng Chen , Kai Liu , Jingkai Wang , Xianglong Yan , Jianze Li , Ziqing Zhang , Jue Gong , Jiatong Li , Lei Sun , Xiaoyang Liu , Radu Timofte , Yulun Zhang , Jihye Park , Yoonjin Im , Hyungju Chun , Hyunhee Park , MinKyu Park , Zheng Xie , Xiangyu Kong , Weijun Yuan , Zhan Li , Qiurong Song , Luen Zhu , Fengkai Zhang , Xinzhe Zhu , Junyang Chen , Congyu Wang , Yixin Yang , Zhaorun Zhou , Jiangxin Dong , Jinshan Pan , Shengwei Wang , Jiajie Ou , Baiang Li , Sizhuo Ma , Qiang Gao , Jusheng Zhang , Jian Wang , Keze Wang , Yijiao Liu , Yingsi Chen , Hui Li , Yu Wang , Congchao Zhu , Saeed Ahmad , Ik Hyun Lee , Jun Young Park , Ji Hwan Yoon , Kainan Yan , Zian Wang , Weibo Wang , Shihao Zou , Chao Dong , Wei Zhou , Linfeng Li , Jaeseong Lee , Jaeho Chae , Jinwoo Kim , Seonjoo Kim , Yucong Hong , Zhenming Yan , Junye Chen , Ruize Han , Song Wang , Yuxuan Jiang , Chengxi Zeng , Tianhao Peng , Fan Zhang , David Bull , Tongyao Mu , Qiong Cao , Yifan Wang , Youwei Pan , Leilei Cao , Xiaoping Peng , Wei Deng , Yifei Chen , Wenbo Xiong , Xian Hu , Yuxin Zhang , Xiaoyun Cheng , Yang Ji , Zonghao Chen , Zhihao Xue , Junqin Hu , Nihal Kumar , Snehal Singh Tomar , Klaus Mueller , Surya Vashisth , Prateek Shaily , Jayant Kumar , Hardik Sharma , Ashish Negi , Sachin Chaudhary , Akshay Dudhane , Praful Hambarde , Amit Shukla , Shijun Shi , Jiangning Zhang , Yong Liu , Kai Hu , Jing Xu , Xianfang Zeng , Amitesh M , Hariharan S , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Nishalini K , Sreenath K A , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Shuling Zheng , Zhiheng Fu , Feng Zhang , Zhanglu Chen , Boyang Yao , Nikhil Pathak , Aagam Jain , Milan Kumar , Kishor Upla , Vivek Chavda , Sarang N S , Raghavendra Ramachandra , Zhipeng Zhang , Qi Wang , Shiyu Wang , Jiachen Tu , Guoyi Xu , Yaoxin Jiang , Jiajia Liu , Yaokun Shi , Yuqi Li , Chuanguang Yang , Weilun Feng , Zhuzhi Hong , Hao Wu , Junming Liu , Yingli Tian , Amish Bhushan Kulkarni , Tejas R R Shet , Saakshi M Vernekar , Nikhil Akalwadi , Kaushik Mallibhat , Ramesh Ashok Tabib , Uma Mudenagudi , Yuwen Pan , Tianrun Chen , Deyi Ji , Qi Zhu , Lanyun Zhu , Heyan Zhangyi

In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration,…

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zhe Cao , Jin Zhang , Ruiheng Zhang

In the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the…

Information Retrieval · Computer Science 2022-11-21 Min Cao , Shiping Li , Juntao Li , Liqiang Nie , Min Zhang

Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most existing benchmarks remain confined to closed-set supervised visual tasks and do not support…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Shuoyuan Wang , Yiran Wang , Hongxin Wei

Progress in a research field can be hard to assess, in particular when many concurrent methods are proposed in a short period of time. This is the case in digital pathology, where many foundation models have been released recently to serve…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Pierre Marza , Leo Fillioux , Sofiène Boutaj , Kunal Mahatha , Christian Desrosiers , Pablo Piantanida , Jose Dolz , Stergios Christodoulidis , Maria Vakalopoulou

While generative AI systems have gained popularity in diverse applications, their potential to produce harmful outputs limits their trustworthiness and usability in different applications. Recent years have seen growing interest in engaging…

Human-Computer Interaction · Computer Science 2025-04-01 Matheus Kunzler Maldaner , Wesley Hanwen Deng , Jason Hong , Ken Holstein , Motahhare Eslami

Vision-Language Models (VLMs) have achieved strong performance on standard vision-language benchmarks, yet often rely on surface-level recognition rather than deeper reasoning. We propose visual word puzzles as a challenging alternative, as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Ali Najar , Alireza Mirrokni , Arshia Izadyari , Sadegh Mohammadian , Amir Homayoon Sharifizade , Asal Meskin , Mobin Bagherian , Ehsaneddin Asgari

Image retrieval with hybrid-modality queries, also known as composing text and image for image retrieval (CTI-IR), is a retrieval task where the search intention is expressed in a more complex query format, involving both vision and text…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Yida Zhao , Yuqing Song , Qin Jin

This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchmark that targets a practical yet difficult HDR imaging setting, where exposure…

We introduce the new Birds-to-Words dataset of 41k sentences describing fine-grained differences between photographs of birds. The language collected is highly detailed, while remaining understandable to the everyday observer (e.g.,…

Computation and Language · Computer Science 2019-11-15 Maxwell Forbes , Christine Kaeser-Chen , Piyush Sharma , Serge Belongie

Measuring advances in retrieval requires test collections with relevance judgments that can faithfully distinguish systems. This paper presents NeuCLIRTech, an evaluation collection for cross-language retrieval over technical information.…

Information Retrieval · Computer Science 2026-02-06 Dawn Lawrie , James Mayfield , Eugene Yang , Andrew Yates , Sean MacAvaney , Ronak Pradeep , Scott Miller , Paul McNamee , Luca Soldaini

We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce visually plausible pictures. To quantify the progress, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Minheng Ni , Zhengyuan Yang , Yaowen Zhang , Linjie Li , Chung-Ching Lin , Kevin Lin , Zhendong Wang , Xiaofei Wang , Shujie Liu , Lei Zhang , Wangmeng Zuo , Lijuan Wang