English

ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models

Computer Vision and Pattern Recognition 2026-04-01 v5 Machine Learning

Abstract

Large Vision-Language Models (LVLMs) excel at captioning, visual question answering, and robotics by combining vision and language, yet they often miss obvious objects or hallucinate nonexistent ones in atypical scenes. We examine these failures through the lens of uncertainty, focusing on contextual incongruity, where objects appear unexpectedly or fail to appear in expected contexts, and show that such cases increase recognition difficulty for state-of-the-art LVLMs. To study this regime, we introduce the Object Recognition in Incongruous Context (ORIC) framework, which constructs incongruous object-context pairs through two complementary strategies: (1) LLM-guided sampling to identify hard-to-recognize objects present in the image and (2) CLIP-guided sampling to mine plausible but absent ones. Applied to MSCOCO, ORIC creates ORIC-Bench and ORIC-style training data. Evaluating 18 LVLMs and 2 open-vocabulary detectors reveals significant degradation and bias under incongruous contexts. Visual Reinforcement Fine-Tuning of Qwen3-VL-8B-Instruct on 600 ORIC samples improves performance on ORIC-Bench, AMBER, and HallusionBench. Overall, we show that contextual incongruity is a key source of uncertainty and provide tools for more reliable LVLMs. The dataset and code are publicly available at https://github.com/ZhaoyangLi-1/ORIC.

Keywords

Cite

@article{arxiv.2509.15695,
  title  = {ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models},
  author = {Zhaoyang Li and Zhan Ling and Yuchen Zhou and Litian Gong and Erdem Bıyık and Hao Su},
  journal= {arXiv preprint arXiv:2509.15695},
  year   = {2026}
}

Comments

We request withdrawal of this paper because one of the listed institutional affiliations was included without proper authorization. This issue cannot be resolved through a simple revision, and we therefore request withdrawal to prevent dissemination of incorrect or unauthorized affiliation information

R2 v1 2026-07-01T05:45:19.236Z