English

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

Computer Vision and Pattern Recognition 2026-07-06 v1

Abstract

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representations. Even when enriched with auxiliary cues, relational interactions remain implicitly encoded within individual anchor features. The resulting visual representation remains flat and unary-only, limiting its ability to align with the structured nature of language. In this work, we propose a Structured Visual Compositional Representation (SVCR) learning framework for WREC. Rather than implicitly encoding relations within unary anchors, the proposed SVCR explicitly models both unary object embeddings and pairwise relational embeddings, forming a structured visual representation space. We further introduce a compositional alignment mechanism that matches unary and pairwise visual representations with their corresponding textual embeddings in a unified manner, enabling compositional visual-textual matching under weak supervision. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg show that the proposed SVCR achieves state-of-the-art performance. These results demonstrate the effectiveness of explicit structured visual representations and visual-textual alignment for WREC.

Cite

@article{arxiv.2607.04638,
  title  = {Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension},
  author = {Lian Xu and Mohammed Bennamoun and Farid Boussaid and Hamid Laga and Yulan Guo and Dan Xu},
  journal= {arXiv preprint arXiv:2607.04638},
  year   = {2026}
}

Comments

Accepted at ECCV 2026