English

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

Computation and Language 2025-11-24 v1

Abstract

We propose a zero-shot method for Natural Language Inference (NLI) that leverages multimodal representations by grounding language in visual contexts. Our approach generates visual representations of premises using text-to-image models and performs inference by comparing these representations with textual hypotheses. We evaluate two inference techniques: cosine similarity and visual question answering. Our method achieves high accuracy without task-specific fine-tuning, demonstrating robustness against textual biases and surface heuristics. Additionally, we design a controlled adversarial dataset to validate the robustness of our approach. Our findings suggest that leveraging visual modality as a meaning representation provides a promising direction for robust natural language understanding.

Keywords

Cite

@article{arxiv.2511.17358,
  title  = {Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding},
  author = {Daniil Ignatev and Ayman Santeer and Albert Gatt and Denis Paperno},
  journal= {arXiv preprint arXiv:2511.17358},
  year   = {2025}
}
R2 v1 2026-07-01T07:48:58.486Z