中文
相关论文

相关论文: NICE: CVPR 2023 Challenge on Zero-shot Image Capti…

200 篇论文

This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Recognition (OCR) models, namely, the recognition of unseen…

计算机视觉与模式识别 · 计算机科学 2022-09-15 Sergi Garcia-Bordils , Andrés Mafla , Ali Furkan Biten , Oren Nuriel , Aviad Aberdam , Shai Mazor , Ron Litman , Dimosthenis Karatzas

We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated news articles, image…

计算机视觉与模式识别 · 计算机科学 2021-09-15 Fuxiao Liu , Yinghan Wang , Tianlu Wang , Vicente Ordonez

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

计算机视觉与模式识别 · 计算机科学 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li

Aesthetic image captioning (AIC) refers to the multi-modal task of generating critical textual feedbacks for photographs. While in natural image captioning (NIC), deep models are trained in an end-to-end manner using large curated datasets…

计算机视觉与模式识别 · 计算机科学 2019-08-30 Koustav Ghosal , Aakanksha Rana , Aljosa Smolic

This paper reviews the NTIRE 2020 challenge on real image denoising with focus on the newly introduced dataset, the proposed methods and their results. The challenge is a new version of the previous NTIRE 2019 challenge on real image…

Image captioning evaluation remains a significant challenge, as vision-language models evolve toward more challenging capabilities such as generating long-form and context-rich descriptions. State-of-the-art evaluation metrics involve…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Gonçalo Gomes , Bruno Martins , Chrysoula Zerva

Unpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Peipei Zhu , Xiao Wang , Lin Zhu , Zhenglong Sun , Weishi Zheng , Yaowei Wang , Changwen Chen

Dataset bias in vision-language tasks is becoming one of the main problems which hinders the progress of our community. Existing solutions lack a principled analysis about why modern image captioners easily collapse into dataset bias. In…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Xu Yang , Hanwang Zhang , Jianfei Cai

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety. In this work, we propose an effective and scalable…

机器人学 · 计算机科学 2025-12-01 Sajjad Pakdamansavoji , Mozhgan Pourkeshavarz , Adam Sigal , Zhiyuan Li , Rui Heng Yang , Amir Rasouli

This paper presents a comprehensive review of the NTIRE 2025 Low-Light Image Enhancement (LLIE) Challenge, highlighting the proposed solutions and final outcomes. The objective of the challenge is to identify effective networks capable of…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xiaoning Liu , Zongwei Wu , Florin-Alexandru Vasluianu , Hailong Yan , Bin Ren , Yulun Zhang , Shuhang Gu , Le Zhang , Ce Zhu , Radu Timofte , Kangbiao Shi , Yixu Feng , Tao Hu , Yu Cao , Peng Wu , Yijin Liang , Yanning Zhang , Qingsen Yan , Han Zhou , Wei Dong , Yan Min , Mohab Kishawy , Jun Chen , Pengpeng Yu , Anjin Park , Seung-Soo Lee , Young-Joon Park , Zixiao Hu , Junyv Liu , Huilin Zhang , Jun Zhang , Fei Wan , Bingxin Xu , Hongzhe Liu , Cheng Xu , Weiguo Pan , Songyin Dai , Xunpeng Yi , Qinglong Yan , Yibing Zhang , Jiayi Ma , Changhui Hu , Kerui Hu , Donghang Jing , Tiesheng Chen , Zhi Jin , Hongjun Wu , Biao Huang , Haitao Ling , Jiahao Wu , Dandan Zhan , G Gyaneshwar Rao , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , Ramesh Ashok Tabib , Uma Mudenagudi , Ruirui Lin , Guoxi Huang , Nantheera Anantrasirichai , Qirui Yang , Alexandru Brateanu , Ciprian Orhei , Cosmin Ancuti , Daniel Feijoo , Juan C. Benito , Álvaro García , Marcos V. Conde , Yang Qin , Raul Balmez , Anas M. Ali , Bilel Benjdira , Wadii Boulila , Tianyi Mao , Huan Zheng , Yanyan Wei , Shengeng Tang , Dan Guo , Zhao Zhang , Sabari Nathan , K Uma , A Sasithradevi , B Sathya Bama , S. Mohamed Mansoor Roomi , Ao Li , Xiangtao Zhang , Zhe Liu , Yijie Tang , Jialong Tang , Zhicheng Fu , Gong Chen , Joe Nasti , John Nicholson , Zeyu Xiao , Zhuoyuan Li , Ashutosh Kulkarni , Prashant W. Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Duan Liu , Weile Li , Hangyuan Lu , Rixian Liu , Tengfeng Wang , Jinxing Liang , Chenxin Yu

Connecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful…

计算机视觉与模式识别 · 计算机科学 2021-12-02 Matteo Stefanini , Marcella Cornia , Lorenzo Baraldi , Silvia Cascianelli , Giuseppe Fiameni , Rita Cucchiara

Zero-shot composed image retrieval (ZS-CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing the desired…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Yongcong Ye , Kai Zhang , Yanghai Zhang , Enhong Chen , Longfei Li , Jun Zhou

Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on paired image-text data. To caption an image, they proceed by textually decoding a text-aligned…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Lorenzo Bianchi , Giacomo Pacini , Fabio Carrara , Nicola Messina , Giuseppe Amato , Fabrizio Falchi

This paper introduces a new benchmark for large-scale image similarity detection. This benchmark is used for the Image Similarity Challenge at NeurIPS'21 (ISC2021). The goal is to determine whether a query image is a modified copy of any…

Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pretrained…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Zequn Zeng , Yan Xie , Hao Zhang , Chiyu Chen , Zhengjue Wang , Bo Chen

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Israa A. Albadarneh , Bassam H. Hammo , Omar S. Al-Kadi

The image captioning task is about to generate suitable descriptions from images. For this task there can be several challenges such as accuracy, fluency and diversity. However there are few metrics that can cover all these properties while…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Chao Zeng , Sam Kwong

The 2021 Image Similarity Challenge introduced a dataset to serve as a new benchmark to evaluate recent image copy detection methods. There were 200 participants to the competition. This paper presents a quantitative and qualitative…