English
Related papers

Related papers: NICE: CVPR 2023 Challenge on Zero-shot Image Capti…

200 papers

This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Recognition (OCR) models, namely, the recognition of unseen…

Computer Vision and Pattern Recognition · Computer Science 2022-09-15 Sergi Garcia-Bordils , Andrés Mafla , Ali Furkan Biten , Oren Nuriel , Aviad Aberdam , Shai Mazor , Ron Litman , Dimosthenis Karatzas

We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated news articles, image…

Computer Vision and Pattern Recognition · Computer Science 2021-09-15 Fuxiao Liu , Yinghan Wang , Tianlu Wang , Vicente Ordonez

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li

Aesthetic image captioning (AIC) refers to the multi-modal task of generating critical textual feedbacks for photographs. While in natural image captioning (NIC), deep models are trained in an end-to-end manner using large curated datasets…

Computer Vision and Pattern Recognition · Computer Science 2019-08-30 Koustav Ghosal , Aakanksha Rana , Aljosa Smolic

This paper reviews the NTIRE 2020 challenge on real image denoising with focus on the newly introduced dataset, the proposed methods and their results. The challenge is a new version of the previous NTIRE 2019 challenge on real image…

Image captioning evaluation remains a significant challenge, as vision-language models evolve toward more challenging capabilities such as generating long-form and context-rich descriptions. State-of-the-art evaluation metrics involve…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Gonçalo Gomes , Bruno Martins , Chrysoula Zerva

Unpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Peipei Zhu , Xiao Wang , Lin Zhu , Zhenglong Sun , Weishi Zheng , Yaowei Wang , Changwen Chen

Dataset bias in vision-language tasks is becoming one of the main problems which hinders the progress of our community. Existing solutions lack a principled analysis about why modern image captioners easily collapse into dataset bias. In…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Xu Yang , Hanwang Zhang , Jianfei Cai

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety. In this work, we propose an effective and scalable…

This paper presents a comprehensive review of the NTIRE 2025 Low-Light Image Enhancement (LLIE) Challenge, highlighting the proposed solutions and final outcomes. The objective of the challenge is to identify effective networks capable of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xiaoning Liu , Zongwei Wu , Florin-Alexandru Vasluianu , Hailong Yan , Bin Ren , Yulun Zhang , Shuhang Gu , Le Zhang , Ce Zhu , Radu Timofte , Kangbiao Shi , Yixu Feng , Tao Hu , Yu Cao , Peng Wu , Yijin Liang , Yanning Zhang , Qingsen Yan , Han Zhou , Wei Dong , Yan Min , Mohab Kishawy , Jun Chen , Pengpeng Yu , Anjin Park , Seung-Soo Lee , Young-Joon Park , Zixiao Hu , Junyv Liu , Huilin Zhang , Jun Zhang , Fei Wan , Bingxin Xu , Hongzhe Liu , Cheng Xu , Weiguo Pan , Songyin Dai , Xunpeng Yi , Qinglong Yan , Yibing Zhang , Jiayi Ma , Changhui Hu , Kerui Hu , Donghang Jing , Tiesheng Chen , Zhi Jin , Hongjun Wu , Biao Huang , Haitao Ling , Jiahao Wu , Dandan Zhan , G Gyaneshwar Rao , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , Ramesh Ashok Tabib , Uma Mudenagudi , Ruirui Lin , Guoxi Huang , Nantheera Anantrasirichai , Qirui Yang , Alexandru Brateanu , Ciprian Orhei , Cosmin Ancuti , Daniel Feijoo , Juan C. Benito , Álvaro García , Marcos V. Conde , Yang Qin , Raul Balmez , Anas M. Ali , Bilel Benjdira , Wadii Boulila , Tianyi Mao , Huan Zheng , Yanyan Wei , Shengeng Tang , Dan Guo , Zhao Zhang , Sabari Nathan , K Uma , A Sasithradevi , B Sathya Bama , S. Mohamed Mansoor Roomi , Ao Li , Xiangtao Zhang , Zhe Liu , Yijie Tang , Jialong Tang , Zhicheng Fu , Gong Chen , Joe Nasti , John Nicholson , Zeyu Xiao , Zhuoyuan Li , Ashutosh Kulkarni , Prashant W. Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Duan Liu , Weile Li , Hangyuan Lu , Rixian Liu , Tengfeng Wang , Jinxing Liang , Chenxin Yu

Connecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Matteo Stefanini , Marcella Cornia , Lorenzo Baraldi , Silvia Cascianelli , Giuseppe Fiameni , Rita Cucchiara

Zero-shot composed image retrieval (ZS-CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing the desired…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yongcong Ye , Kai Zhang , Yanghai Zhang , Enhong Chen , Longfei Li , Jun Zhou

This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE workshop at CVPR 2026. The goal of this challenge was to develop detection models capable of…

Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on paired image-text data. To caption an image, they proceed by textually decoding a text-aligned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Lorenzo Bianchi , Giacomo Pacini , Fabio Carrara , Nicola Messina , Giuseppe Amato , Fabrizio Falchi

This paper introduces a new benchmark for large-scale image similarity detection. This benchmark is used for the Image Similarity Challenge at NeurIPS'21 (ISC2021). The goal is to determine whether a query image is a modified copy of any…

Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pretrained…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Zequn Zeng , Yan Xie , Hao Zhang , Chiyu Chen , Zhengjue Wang , Bo Chen

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Israa A. Albadarneh , Bassam H. Hammo , Omar S. Al-Kadi

The image captioning task is about to generate suitable descriptions from images. For this task there can be several challenges such as accuracy, fluency and diversity. However there are few metrics that can cover all these properties while…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Chao Zeng , Sam Kwong

The 2021 Image Similarity Challenge introduced a dataset to serve as a new benchmark to evaluate recent image copy detection methods. There were 200 participants to the competition. This paper presents a quantitative and qualitative…