English
Related papers

Related papers: MUNIChus: Multilingual News Image Captioning Bench…

200 papers

Image captioning is a multimodal task involving computer vision and natural language processing, where the goal is to learn a mapping from the image to its natural language description. In general, the mapping function is learned from a…

Computer Vision and Pattern Recognition · Computer Science 2018-07-19 Jiuxiang Gu , Shafiq Joty , Jianfei Cai , Gang Wang

Image captioning has so far been explored mostly in English, as most available datasets are in this language. However, the application of image captioning should not be restricted by language. Only few studies have been conducted for image…

Computation and Language · Computer Science 2017-08-16 Weiyu Lan , Xirong Li , Jianfeng Dong

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text…

Computation and Language · Computer Science 2025-09-29 Xiaolong Wang , Zhaolu Kang , Wangyuxuan Zhai , Xinyue Lou , Yunghwei Lai , Ziyue Wang , Yawen Wang , Kaiyu Huang , Yile Wang , Peng Li , Yang Liu

This research explores the realm of neural image captioning using deep learning models. The study investigates the performance of different neural architecture configurations, focusing on the inject architecture, and proposes a novel…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Pooja Bhatnagar , Sai Mrunaal , Sachin Kamnure

Automatically generating descriptive captions for images is a well-researched area in computer vision. However, existing evaluation approaches focus on measuring the similarity between two sentences disregarding fine-grained semantics of…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Philipp Harzig , Dan Zecha , Rainer Lienhart , Carolin Kaiser , René Schallner

As computer-generated content and deepfakes make steady improvements, semantic approaches to multimedia forensics will become more important. In this paper, we introduce a novel classification architecture for identifying semantic…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Scott McCrae , Kehan Wang , Avideh Zakhor

Automated image captioning using the content from the image is very appealing when done by harnessing the capability of computer vision and natural language processing. Extensive research has been done in the field with a major focus on the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Wasim Akram Khan , Anil Kumar Vuppala

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li

We present a large, multilingual study into how vision constrains linguistic choice, covering four languages and five linguistic properties, such as verb transitivity or use of numerals. We propose a novel method that leverages existing…

Computation and Language · Computer Science 2023-02-10 Uri Berger , Lea Frermann , Gabriel Stanovsky , Omri Abend

Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zitong Xu , Huiyu Duan , Shengyao Qin , Guangyu Yang , Guangji Ma , Xiongkuo Min , Ke Gu , Guangtao Zhai , Patrick Le Callet

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Harsh Agrawal , Karan Desai , Yufei Wang , Xinlei Chen , Rishabh Jain , Mark Johnson , Dhruv Batra , Devi Parikh , Stefan Lee , Peter Anderson

Generating image descriptions in different languages is essential to satisfy users worldwide. However, it is prohibitively expensive to collect large-scale paired image-caption dataset for every target language which is critical for…

Computer Vision and Pattern Recognition · Computer Science 2019-08-16 Yuqing Song , Shizhe Chen , Yida Zhao , Qin Jin

Deep neural networks have achieved great successes on the image captioning task. However, most of the existing models depend heavily on paired image-sentence datasets, which are very expensive to acquire. In this paper, we make the first…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Yang Feng , Lin Ma , Wei Liu , Jiebo Luo

Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present a challenging setting where the caption requires such named…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Hammad A. Ayyubi , Tianqi Liu , Arsha Nagrani , Xudong Lin , Mingda Zhang , Anurag Arnab , Feng Han , Yukun Zhu , Jialu Liu , Shih-Fu Chang

Recent advancements in information retrieval have highlighted the potential of integrating visual and textual information, yet effective reranking for image-text documents remains challenging due to the modality gap and scarcity of aligned…

Information Retrieval · Computer Science 2026-01-29 Hongyi Cai

Unsupervised image-to-image translation is an important and challenging problem in computer vision. Given an image in the source domain, the goal is to learn the conditional distribution of corresponding images in the target domain, without…

Computer Vision and Pattern Recognition · Computer Science 2018-08-16 Xun Huang , Ming-Yu Liu , Serge Belongie , Jan Kautz

The widespread dissemination of false information through manipulative tactics that combine deceptive text and images threatens the integrity of reliable sources of information. While there has been research on detecting fake news in high…

Computation and Language · Computer Science 2024-10-15 Shubhi Bansal , Nishit Sushil Singh , Shahid Shafi Dar , Nagendra Kumar

Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate caption and a set of…

Computation and Language · Computer Science 2019-12-20 Huiyuan Xie , Tom Sherborne , Alexander Kuhnle , Ann Copestake

Recently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images is often limited to just picking the answer from a…

In this paper we describe a novel framework and algorithms for discovering image patch patterns from a large corpus of weakly supervised image-caption pairs generated from news events. Current pattern mining techniques attempt to find…

Computer Vision and Pattern Recognition · Computer Science 2016-01-06 Hongzhi Li , Joseph G. Ellis , Shih-Fu Chang