English
Related papers

Related papers: Improving Multimodal Datasets with Image Captionin…

200 papers

Recent advancements in text-to-image generation using diffusion models have significantly improved the quality of generated images and expanded the ability to depict a wide range of objects. However, ensuring that these models adhere…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Michail Tarasiou , Stylianos Moschoglou , Jiankang Deng , Stefanos Zafeiriou

While deep-learning models have been shown to perform well on image-to-text datasets, it is difficult to use them in practice for captioning images. This is because captions traditionally tend to be context-dependent and offer complementary…

Machine Learning · Computer Science 2023-06-07 Shinjini Ghosh , Sagnik Anupam

Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice. We argue that to close this gap, it is vital to distinguish descriptions from captions…

Computation and Language · Computer Science 2022-10-31 Elisa Kreiss , Fei Fang , Noah D. Goodman , Christopher Potts

Visual recognition in a low-data regime is challenging and often prone to overfitting. To mitigate this issue, several data augmentation strategies have been proposed. However, standard transformations, e.g., rotation, cropping, and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Aniket Roy , Anshul Shah , Ketul Shah , Anirban Roy , Rama Chellappa

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning approaches heavily rely on high-quality image-caption pairs, which…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Chuanyang Jin

This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary samples from the latent morphospace of a species, each…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ziheng Zhang , Xinyue Ma , Arpita Chowdhury , Elizabeth G. Campolongo , Matthew J. Thompson , Net Zhang , Samuel Stevens , Hilmar Lapp , Tanya Berger-Wolf , Yu Su , Wei-Lun Chao , Jianyang Gu

Neural captioners are typically trained to mimic human-generated references without optimizing for any specific communication goal, leading to problems such as the generation of vague captions. In this paper, we show that fine-tuning an…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Roberto Dessì , Michele Bevilacqua , Eleonora Gualdoni , Nathanael Carraz Rakotonirina , Francesca Franzon , Marco Baroni

In this study, we introduce a novel cover image generation task that produces both a concise summary and a visually corresponding image from a given text-only document. Because no existing datasets are available for this task, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hyeyeon Kim , Sungwoo Han , Jingun Kwon , Hidetaka Kamigaito , Manabu Okumura

Automated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and even unsafe model predictions. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Lu Yu , Malvina Nikandrou , Jiali Jin , Verena Rieser

Currently, image-text-driven multi-modal deep learning models have demonstrated their outstanding potential in many fields. In practice, tasks centered around facial images have broad application prospects. This paper presents…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Dawei Dai , YuTang Li , YingGe Liu , Mingming Jia , Zhang YuanHui , Guoyin Wang

Many top-performing image captioning models rely solely on object features computed with an object detection model to generate image descriptions. However, recent studies propose to directly use scene graphs to introduce information about…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Victor Milewski , Marie-Francine Moens , Iacer Calixto

The availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is…

Computer Vision and Pattern Recognition · Computer Science 2019-03-04 Yazhou Yao , Jian Zhang , Fumin Shen , Li Liu , Fan Zhu , Dongxiang Zhang , Heng-Tao Shen

Aesthetic image captioning (AIC) refers to the multi-modal task of generating critical textual feedbacks for photographs. While in natural image captioning (NIC), deep models are trained in an end-to-end manner using large curated datasets…

Computer Vision and Pattern Recognition · Computer Science 2019-08-30 Koustav Ghosal , Aakanksha Rana , Aljosa Smolic

Benefiting from advances in machine vision and natural language processing techniques, current image captioning systems are able to generate detailed visual descriptions. For the most part, these descriptions represent an objective…

Computer Vision and Pattern Recognition · Computer Science 2020-04-16 Omid Mohamad Nezami , Mark Dras , Stephen Wan , Cecile Paris

Web-crawled datasets have enabled remarkable generalization capabilities in recent image-text models such as CLIP (Contrastive Language-Image pre-training) or Flamingo, but little is known about the dataset creation processes. In this work,…

Machine Learning · Computer Science 2023-02-02 Thao Nguyen , Gabriel Ilharco , Mitchell Wortsman , Sewoong Oh , Ludwig Schmidt

The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods. Our work studies this question…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Shibani Santurkar , Yann Dubois , Rohan Taori , Percy Liang , Tatsunori Hashimoto

Deep neural networks have achieved great successes on the image captioning task. However, most of the existing models depend heavily on paired image-sentence datasets, which are very expensive to acquire. In this paper, we make the first…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Yang Feng , Lin Ma , Wei Liu , Jiebo Luo

Curation methods for massive vision-language datasets trade off between dataset size and quality. However, even the highest quality of available curated captions are far too short to capture the rich visual detail in an image. To show the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jack Urbanek , Florian Bordes , Pietro Astolfi , Mary Williamson , Vasu Sharma , Adriana Romero-Soriano

Recent advances in image captioning have focused on scaling the data and model size, substantially increasing the cost of pre-training and finetuning. As an alternative to large models, we present SmallCap, which generates a caption…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Rita Ramos , Bruno Martins , Desmond Elliott , Yova Kementchedjhieva

We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synthetic datasets of images and corresponding captions at scale,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Hasan Abed Al Kader Hammoud , Hani Itani , Fabio Pizzati , Philip Torr , Adel Bibi , Bernard Ghanem