English

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

Artificial Intelligence 2026-07-20 v1

Abstract

Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this as a fundamental representation learning problem, providing the first evidence that negation does not form a linearly or non-linearly separable class in the latent spaces of standard vision-language models (VLMs). We demonstrate that pretrained embeddings primarily encode modality-specific features, lacking a generalizable negation signal. To overcome this, we propose a novel cross-modal attention architecture that explicitly models inter-modal dependencies, achieving performance gains of up to +7.03% F1 over unimodal baselines. Our analysis reveals a key asymmetry: while textual negation often appears independently, visual negation is semantically dependent on linguistic context, a finding validated through our statistical analysis of 3,222 political video-text pairs automatically annotated via \textsc{Qwen2.5-VL}. By combining this analysis with self-supervised video representations (JEPA2), we advance the modeling of temporal negation. This work provides new methods and insights for learning robust, semantically-aligned representations in multimodal systems.

Keywords

Cite

@article{arxiv.2607.17712,
  title  = {Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution},
  author = {Ali AbuSaleh and Leon Hammerla and Alexander Mehler},
  journal= {arXiv preprint arXiv:2607.17712},
  year   = {2026}
}

Comments

This manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861