English
Related papers

Related papers: Imagery as Inquiry: Exploring A Multimodal Dataset…

200 papers

In this paper, we introduce an attribute-based interactive image search which can leverage human-in-the-loop feedback to iteratively refine image search results. We study active image search where human feedback is solicited exclusively in…

Computer Vision and Pattern Recognition · Computer Science 2018-09-25 Bryan A. Plummer , M. Hadi Kiapour , Shuai Zheng , Robinson Piramuthu

We introduce the Multi30K dataset to stimulate multilingual multimodal research. Recent advances in image description have been demonstrated on English-language datasets almost exclusively, but image description should not be limited to…

Computation and Language · Computer Science 2016-05-03 Desmond Elliott , Stella Frank , Khalil Sima'an , Lucia Specia

Is aesthetic impact different from beauty? Is visual salience a reflection of its capacity for effective communication? We present Impressions, a novel dataset through which to investigate the semiotics of images, and how specific visual…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Julia Kruk , Caleb Ziems , Diyi Yang

Recently, numbers of works shows that the performance of neural machine translation (NMT) can be improved to a certain extent with using visual information. However, most of these conclusions are drawn from the analysis of experimental…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 ZhenHao Tang , XiaoBing Zhang , Zi Long , XiangHua Fu

Deep neural networks have achieved promising results in automatic image captioning due to their effective representation learning and context-based content generation capabilities. As a prominent type of deep features used in many of the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Ali Abedi , Hossein Karshenas , Peyman Adibi

Humans express feelings or emotions via different channels. Take language as an example, it entails different sentiments under different visual-acoustic contexts. To precisely understand human intentions as well as reduce the…

Artificial Intelligence · Computer Science 2021-11-17 Ting Wu , Junjie Peng , Wenqiang Zhang , Huiran Zhang , Chuanshuai Ma , Yansong Huang

Due to the availability of increasingly large amounts of visual data, there is a growing need for tools that can help users find relevant images. While existing tools can perform image retrieval based on similarity or metadata, they fall…

Human-Computer Interaction · Computer Science 2024-01-22 Celeste Barnaby , Qiaochu Chen , Chenglong Wang , Isil Dillig

Multimodal sentiment analysis has a wide range of applications due to its information complementarity in multimodal interactions. Previous works focus more on investigating efficient joint representations, but they rarely consider the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Rongfei Chen , Wenju Zhou , Yang Li , Huiyu Zhou

Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the images can be repetitive. This paper proposes a Prompt Expansion framework that helps users generate…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Siddhartha Datta , Alexander Ku , Deepak Ramachandran , Peter Anderson

Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning (e.g., visual question answering and image retrieval). We are interested in shedding light on the quality of their…

Computation and Language · Computer Science 2021-06-18 Lisa Anne Hendricks , Aida Nematzadeh

Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how…

Information Retrieval · Computer Science 2025-08-08 Hongyu Zhou , Yinan Zhang , Aixin Sun , Zhiqi Shen

A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations. In this paper, we use logic-based representations as unified meaning…

Computation and Language · Computer Science 2019-06-11 Riko Suzuki , Hitomi Yanaka , Masashi Yoshikawa , Koji Mineshima , Daisuke Bekki

Multimodal entity linking (MEL), a task aimed at linking mentions within multimodal contexts to their corresponding entities in a knowledge base (KB), has attracted much attention due to its wide applications in recent years. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Hongze Mi , Jinyuan Li , Xuying Zhang , Haoran Cheng , Jiahao Wang , Di Sun , Gang Pan

Humans describe images in terms of nouns and adjectives while algorithms operate on images represented as sets of pixels. Bridging this gap between how humans would like to access images versus their typical representation is the goal of…

In recent years, text-guided image manipulation has gained increasing attention in the multimedia and computer vision community. The input to conditional image generation has evolved from image-only to multimodality. In this paper, we study…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Tianhao Zhang , Hung-Yu Tseng , Lu Jiang , Weilong Yang , Honglak Lee , Irfan Essa

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has focused primarily on recognition tasks such as meal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Sabab Ishraq , Aarushi Aarushi , Juncai Jiang , Chen Chen

Facilitated by deep neural networks, video recommendation systems have made significant advances. Existing video recommendation systems directly exploit features from different modalities (e.g., user personal data, user behavior data, video…

Information Retrieval · Computer Science 2020-10-27 Shi Pu , Yijiang He , Zheng Li , Mao Zheng

This paper addresses the generation of explanations with visual examples. Given an input sample, we build a system that not only classifies it to a specific category, but also outputs linguistic explanations and a set of visual examples…

Computer Vision and Pattern Recognition · Computer Science 2019-05-21 Atsushi Kanehira , Tatsuya Harada

Vision-language models can assess visual context in an image and generate descriptive text. While the generated text may be accurate and syntactically correct, it is often overly general. To address this, recent work has used optical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Wes Robbins , Zanyar Zohourianshahzadi , Jugal Kalita

The combination of visual and textual representations has produced excellent results in tasks such as image captioning and visual question answering, but the inference capabilities of multimodal representations are largely untested. In the…

Computation and Language · Computer Science 2020-04-07 Oier Lopez de Lacalle , Ander Salaberria , Aitor Soroa , Gorka Azkune , Eneko Agirre