English

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

Computer Vision and Pattern Recognition 2026-07-14 v1

Abstract

Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-20K, a large-scale triplet-based training set automatically constructed from structured visual metadata through a multi-expert collaborative pipeline, covering counting, depth, distance, and spatial relations; (ii) TriSR, a structured reward framework that jointly supervises attribute grounding, judgment alignment, and triplet consistency under GRPO optimization; and (iii) CoRe-Bench, the first benchmark dedicated to fine-grained cross-image comparative reasoning. Experiments show that CoRe substantially outperforms existing VLMs on CoRe-Bench while remaining competitive on standard multimodal benchmarks, achieving a 28.2-point gain in partial accuracy over the strongest baseline.

Cite

@article{arxiv.2607.12786,
  title  = {CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models},
  author = {Lin Peng and Cong Wan and Zeyu Guo and SongLin Dong and Yihong Gong},
  journal= {arXiv preprint arXiv:2607.12786},
  year   = {2026}
}

Comments

Accepted by ACMMM2026