English
Related papers

Related papers: Multi30K: Multilingual English-German Image Descri…

200 papers

The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning…

Computation and Language · Computer Science 2025-02-18 Gonçalo Gomes , Chrysoula Zerva , Bruno Martins

We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universal representations…

Computation and Language · Computer Science 2021-04-02 Minheng Ni , Haoyang Huang , Lin Su , Edward Cui , Taroon Bharti , Lijuan Wang , Jianfeng Gao , Dongdong Zhang , Nan Duan

Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only…

Recent work has highlighted the advantage of jointly learning grounded sentence representations from multiple languages. However, the data used in these studies has been limited to an aligned scenario: the same images annotated with…

Computation and Language · Computer Science 2019-11-12 Ákos Kádár , Grzegorz Chrupała , Afra Alishahi , Desmond Elliott

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

Computer Vision and Pattern Recognition · Computer Science 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingual image captioning…

Computation and Language · Computer Science 2025-10-21 Mardiyyah Oduwole , Prince Mireku , Fatimo Adebanjo , Oluwatosin Olajide , Mahi Aminu Aliyu , Jekaterina Novikova

Robust automatic fact-checking systems have the potential to combat online misinformation at scale. However, most existing research primarily focuses on English. In this paper, we introduce MultiSynFact, the first large-scale multilingual…

Computation and Language · Computer Science 2025-02-24 Yi-Ling Chung , Aurora Cobo , Pablo Serna

In this paper, a self-guiding multimodal LSTM (sg-LSTM) image captioning model is proposed to handle uncontrolled imbalanced real-world image-sentence dataset. We collect FlickrNYC dataset from Flickr as our testbed with 306,165 images and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-18 Yang Xian , Yingli Tian

We develop ImageNet-Think, a multimodal reasoning dataset designed to aid the development of Vision Language Models (VLMs) with explicit reasoning capabilities. Our dataset is built on 250,000 images from ImageNet21k dataset, providing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Krishna Teja Chitty-Venkata , Murali Emani

Recent integration of Natural Language Processing (NLP) and multimodal models has advanced the field of sports analytics. This survey presents a comprehensive review of the datasets and applications driving these innovations post-2020. We…

Computation and Language · Computer Science 2024-06-19 Haotian Xia , Zhengbang Yang , Yun Zhao , Yuqing Wang , Jingxi Li , Rhys Tracy , Zhuangdi Zhu , Yuan-fang Wang , Hanjie Chen , Weining Shen

State-of-the-art machine translation (MT) models do not use knowledge of any single language's structure; this is the equivalent of asking someone to translate from English to German while knowing neither language. BALM is a framework…

Computation and Language · Computer Science 2019-09-04 Jeffrey Cheng , Chris Callison-Burch

Multimodal named entity recognition (MNER) requires to bridge the gap between language understanding and visual context. While many multimodal neural techniques have been proposed to incorporate images into the MNER task, the model's…

Computation and Language · Computer Science 2021-09-21 Shuguang Chen , Gustavo Aguilar , Leonardo Neves , Thamar Solorio

Automatically creating the description of an image using any natural languages sentence like English is a very challenging task. It requires expertise of both image processing as well as natural language processing. This paper discuss about…

Computer Vision and Pattern Recognition · Computer Science 2018-10-03 Parth Shah , Vishvajit Bakarola , Supriya Pati

This technical report provides extra details of the deep multimodal similarity model (DMSM) which was proposed in (Fang et al. 2015, arXiv:1411.4952). The model is trained via maximizing global semantic similarity between images and their…

Computer Vision and Pattern Recognition · Computer Science 2015-04-29 Xiaodong He , Rupesh Srivastava , Jianfeng Gao , Li Deng

We introduce multi-modal, attention-based neural machine translation (NMT) models which incorporate visual features into different parts of both the encoder and the decoder. We utilise global image features extracted using a pre-trained…

Computation and Language · Computer Science 2017-01-24 Iacer Calixto , Qun Liu , Nick Campbell

Due to the availability of increasingly large amounts of visual data, there is a growing need for tools that can help users find relevant images. While existing tools can perform image retrieval based on similarity or metadata, they fall…

Human-Computer Interaction · Computer Science 2024-01-22 Celeste Barnaby , Qiaochu Chen , Chenglong Wang , Isil Dillig

Visual Genome is a dataset connecting structured image information with English language. We present ``Hindi Visual Genome'', a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task…

Computation and Language · Computer Science 2019-07-23 Shantipriya Parida , Ondřej Bojar , Satya Ranjan Dash

We address the task of text translation on the How2 dataset using a state of the art transformer-based multimodal approach. The question we ask ourselves is whether visual features can support the translation process, in particular, given…

Computation and Language · Computer Science 2019-08-20 Zixiu Wu , Julia Ive , Josiah Wang , Pranava Madhyastha , Lucia Specia

Image captioning is the process of generating a natural language description of an image. Most current image captioning models, however, do not take into account the emotional aspect of an image, which is very relevant to activities and…

Computer Vision and Pattern Recognition · Computer Science 2019-01-28 Omid Mohamad Nezami , Mark Dras , Peter Anderson , Len Hamey

Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is…

Computation and Language · Computer Science 2020-11-13 Douwe Kiela , Suvrat Bhooshan , Hamed Firooz , Ethan Perez , Davide Testuggine