English
Related papers

Related papers: Can visual language models resolve textual ambigui…

200 papers

Visual Word Sense Disambiguation (VWSD) is a novel challenging task with the goal of retrieving an image among a set of candidates, which better represents the meaning of an ambiguous word within a given context. In this paper, we make a…

Computation and Language · Computer Science 2024-04-23 Anastasia Kritharoula , Maria Lymperaiou , Giorgos Stamou

In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. Aligning cross-modal semantics is claimed to be one of the essential capabilities of VLP…

Computation and Language · Computer Science 2022-10-19 Zheng Ma , Shi Zong , Mianzhi Pan , Jianbing Zhang , Shujian Huang , Xinyu Dai , Jiajun Chen

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence…

Computation and Language · Computer Science 2025-08-21 Admitos Passadakis , Yingjin Song , Albert Gatt

Recent studies show that deep vision-only and language-only models--trained on disjoint modalities--nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of where in each network…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Zoe Wanying He , Sean Trott , Meenakshi Khosla

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental image context in…

Computation and Language · Computer Science 2026-05-27 Yifan Jiang , Ruoxi Ning , Sheng Yao , Freda Shi

Capturing semantic relations between sentences, such as entailment, is a long-standing challenge for computational semantics. Logic-based models analyse entailment in terms of possible worlds (interpretations, or situations) where a premise…

Multilingual (or cross-lingual) embeddings represent several languages in a unique vector space. Using a common embedding space enables for a shared semantic between words from different languages. In this paper, we propose to embed images…

Computer Vision and Pattern Recognition · Computer Science 2019-05-15 Maxime Portaz , Hicham Randrianarivo , Adrien Nivaggioli , Estelle Maudet , Christophe Servan , Sylvain Peyronnet

The ideal form of Visual Question Answering requires understanding, grounding and reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most existing VQA benchmarks are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Kang Chen , Xiangqian Wu

Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this work, we focus on investigating the advantages and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Xuejing Liu , Wei Tang , Xinzhe Ni , Jinghui Lu , Rui Zhao , Zechao Li , Fei Tan

The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning…

Computation and Language · Computer Science 2025-02-18 Gonçalo Gomes , Chrysoula Zerva , Bruno Martins

Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Peng Jin , Ryuichi Takanobu , Wancai Zhang , Xiaochun Cao , Li Yuan

Reasoning is central to human intelligence, enabling structured problem-solving across diverse tasks. Recent advances in large language models (LLMs) have greatly enhanced their reasoning abilities in arithmetic, commonsense, and symbolic…

Effectiveness and interpretability are two essential properties for trustworthy AI systems. Most recent studies in visual reasoning are dedicated to improving the accuracy of predicted answers, and less attention is paid to explaining the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Shi Chen , Qi Zhao

Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However, most of the existing MLLMs mainly adopt vision encoders pretrained on coarsely aligned image-text pairs,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Gongwei Chen , Leyang Shen , Rui Shao , Xiang Deng , Liqiang Nie

Large neural networks can now generate jokes, but do they really "understand" humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption,…

Computation and Language · Computer Science 2023-07-07 Jack Hessel , Ana Marasović , Jena D. Hwang , Lillian Lee , Jeff Da , Rowan Zellers , Robert Mankoff , Yejin Choi

Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent…

Computation and Language · Computer Science 2026-02-25 Dilara Torunoğlu-Selamet , Dogukan Arslan , Rodrigo Wilkens , Wei He , Doruk Eryiğit , Thomas Pickard , Adriana S. Pagano , Aline Villavicencio , Gülşen Eryiğit , Ágnes Abuczki , Aida Cardoso , Alesia Lazarenka , Dina Almassova , Amalia Mendes , Anna Kanellopoulou , Antoni Brosa-Rodríguez , Baiba Saulite , Beata Wojtowicz , Bolette Pedersen , Carlos Manuel Hidalgo-Ternero , Chaya Liebeskind , Danka Jokić , Diego Alves , Eleni Triantafyllidi , Erik Velldal , Fred Philippy , Giedre Valunaite Oleskeviciene , Ieva Rizgeliene , Inguna Skadina , Irina Lobzhanidze , Isabell Stinessen Haugen , Jauza Akbar Krito , Jelena M. Marković , Johanna Monti , Josue Alejandro Sauca , Kaja Dobrovoljc , Kingsley O. Ugwuanyi , Laura Rituma , Lilja Øvrelid , Maha Tufail Agro , Manzura Abjalova , Maria Chatzigrigoriou , María del Mar Sánchez Ramos , Marija Pendevska , Masoumeh Seyyedrezaei , Mehrnoush Shamsfard , Momina Ahsan , Muhammad Ahsan Riaz Khan , Nathalie Carmen Hau Norman , Nilay Erdem Ayyıldız , Nina Hosseini-Kivanani , Noémi Ligeti-Nagy , Numaan Naeem , Olha Kanishcheva , Olha Yatsyshyna , Daniil Orel , Petra Giommarelli , Petya Osenova , Radovan Garabik , Regina E. Semou , Rozane Rebechi , Salsabila Zahirah Pranida , Samia Touileb , Sanni Nimb , Sarfraz Ahmad , Sarvinoz Sharipova , Shahar Golan , Shaoxiong Ji , Sopuruchi Christian Aboh , Srdjan Sucur , Stella Markantonatou , Sussi Olsen , Vahide Tajalli , Veronika Lipp , Voula Giouli , Yelda Yeşildal Eraydın , Zahra Saaberi , Zhuohan Xie

Unlike traditional vision-only models, vision language models (VLMs) offer an intuitive way to access visual content through language prompting by combining a large language model (LLM) with a vision encoder. However, both the LLM and the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Paul Gavrikov , Jovita Lukasik , Steffen Jung , Robert Geirhos , M. Jehanzeb Mirza , Margret Keuper , Janis Keuper

Explanations for computer vision models are important tools for interpreting how the underlying models work. However, they are often presented in static formats, which pose challenges for users, including information overload, a gap between…

Human-Computer Interaction · Computer Science 2025-04-16 Indu Panigrahi , Sunnie S. Y. Kim , Amna Liaqat , Rohan Jinturkar , Olga Russakovsky , Ruth Fong , Parastoo Abtahi

Two modalities are often used to convey information in a complementary and beneficial manner, e.g., in online news, videos, educational resources, or scientific publications. The automatic understanding of semantic correlations between text…

Multimedia · Computer Science 2019-06-21 Christian Otto , Matthias Springstein , Avishek Anand , Ralph Ewerth

Figurative language is a challenge for language models since its interpretation is based on the use of words in a way that deviates from their conventional order and meaning. Yet, humans can easily understand and interpret metaphors,…

Computation and Language · Computer Science 2023-06-16 Philipp Wicke
‹ Prev 1 3 4 5 6 7 10 Next ›