English
Related papers

Related papers: The Solution for the ICCV 2023 1st Scientific Figu…

200 papers

Image captioning is a fundamental task in vision-language understanding, where the model predicts a textual informative caption to a given input image. In this paper, we present a simple approach to address this task. We use CLIP encoding…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Ron Mokady , Amir Hertz , Amit H. Bermano

Understanding long text is of great demands in practice but beyond the reach of most language-image pre-training (LIP) models. In this work, we empirically confirm that the key reason causing such an issue is that the training images are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Wei Wu , Kecheng Zheng , Shuailei Ma , Fan Lu , Yuxin Guo , Yifei Zhang , Wei Chen , Qingpei Guo , Yujun Shen , Zheng-Jun Zha

With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX dataset, now including…

Computation and Language · Computer Science 2025-09-01 João Guilherme Alves Santos , Giovana Kerche Bonás , Thales Sales Almeida

PDF documents contain critical visual elements such as figures, tables, and forms whose accurate extraction is essential for document understanding and multimodal retrieval-augmented generation (RAG). Existing PDF parsers often miss complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Meizhu Liu , Yassi Abbasi , Matthew Rowe , Michael Avendi , Paul Li

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Chunlin Zhong , Qiuxia Hou , Zhangjun Zhou , Shuang Hao , Haonan Lu , Yanhao Zhang , He Tang , Xiang Bai

In this paper, we aim to understand whether current language and vision (LaVi) models truly grasp the interaction between the two modalities. To this end, we propose an extension of the MSCOCO dataset, FOIL-COCO, which associates images…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Ravi Shekhar , Sandro Pezzelle , Yauhen Klimovich , Aurelie Herbelot , Moin Nabi , Enver Sangineto , Raffaella Bernardi

We argue that generative text-to-image models often struggle with prompt adherence due to the noisy and unstructured nature of large-scale datasets like LAION-5B. This forces users to rely heavily on prompt engineering to elicit desirable…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Nicholas Merchant , Haitz Sáez de Ocáriz Borde , Andrei Cristian Popescu , Carlos Garcia Jurado Suarez

The conventional training approach for image captioning involves pre-training a network using teacher forcing and subsequent fine-tuning with Self-Critical Sequence Training to maximize hand-crafted captioning metrics. However, when…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Nicholas Moratelli , Davide Caffagni , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

We establish THumB, a rubric-based human evaluation protocol for image captioning models. Our scoring rubrics and their definitions are carefully developed based on machine- and human-generated captions on the MSCOCO dataset. Each caption…

Computation and Language · Computer Science 2022-05-20 Jungo Kasai , Keisuke Sakaguchi , Lavinia Dunagan , Jacob Morrison , Ronan Le Bras , Yejin Choi , Noah A. Smith

Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zitong Xu , Huiyu Duan , Shengyao Qin , Guangyu Yang , Guangji Ma , Xiongkuo Min , Ke Gu , Guangtao Zhai , Patrick Le Callet

This report presents our submission to the MS COCO Captioning Challenge 2015. The method uses Convolutional Neural Network activations as an embedding to find semantically similar images. From these images, the most typical caption is…

Computer Vision and Pattern Recognition · Computer Science 2015-06-15 Martin Kolář , Michal Hradiš , Pavel Zemčík

Automated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and even unsafe model predictions. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Lu Yu , Malvina Nikandrou , Jiali Jin , Verena Rieser

In the domain of vision-language integration, generating detailed image captions poses a significant challenge due to the lack of curated and rich datasets. This study introduces PixLore, a novel method that leverages Querying Transformers…

We introduce a new large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering. First, we identify for 39,181 images taken by people who are blind…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Tai-Yin Chiu , Yinan Zhao , Danna Gurari

Purpose: Our study presents an enhanced approach to medical image caption generation by integrating concept detection into attention mechanisms. Method: This method utilizes sophisticated models to identify critical concepts within medical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Nhi Ngoc-Yen Nguyen , Le-Huy Tu , Dieu-Phuong Nguyen , Nhat-Tan Do , Minh Triet Thai , Bao-Thien Nguyen-Tat

This paper addresses the problem of generating table captions for scholarly documents, which often require additional information outside the table. To this end, we propose a method of retrieving relevant sentences from the paper body, and…

Computation and Language · Computer Science 2021-08-19 Junjie H. Xu , Kohei Shinden , Makoto P. Kato

We investigate the incorporation of visual relationships into the task of supervised image caption generation by proposing a model that leverages detected objects and auto-generated visual relationships to describe images in natural…

Computer Vision and Pattern Recognition · Computer Science 2021-09-24 Maximilian Mozes , Martin Schmitt , Vladimir Golkov , Hinrich Schütze , Daniel Cremers

A wide range of image captioning models has been developed, achieving significant improvement based on popular metrics, such as BLEU, CIDEr, and SPICE. However, although the generated captions can accurately describe the image, they are…

Computer Vision and Pattern Recognition · Computer Science 2020-09-30 Jiuniu Wang , Wenjia Xu , Qingzhong Wang , Antoni B. Chan

The application of Vision-language foundation models (VLFMs) to remote sensing (RS) imagery has garnered significant attention due to their superior capability in various downstream tasks. A key challenge lies in the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yiguo He , Junjie Zhu , Yiying Li , Xiaoyu Zhang , Chunping Qiu , Jun Wang , Qiangjuan Huang , Ke Yang

The task of image captioning aims to generate captions directly from images via the automatically learned cross-modal generator. To build a well-performing generator, existing approaches usually need a large number of described images,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Yang Yang , Hongchen Wei , Hengshu Zhu , Dianhai Yu , Hui Xiong , Jian Yang
‹ Prev 1 4 5 6 7 8 10 Next ›