English

Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models

Artificial Intelligence 2026-05-12 v2 Machine Learning

Abstract

Despite the advanced capabilities of Large Vision-Language Models (LVLMs), they frequently suffer from object hallucination. One reason is that visual features and pretrained textual representations often become intertwined in the deeper network layers. To address this, we propose REVIS, a training-free framework designed to explicitly re-activate this suppressed visual information. Rooted in latent space geometry, REVIS extracts the pure visual information vector via orthogonal projection and employs a calibrated strategy to perform sparse intervention only at the precise depth where suppression occurs. This surgical approach effectively restores visual information with minimal computational cost. Empirical evaluations on standard benchmarks demonstrate that REVIS reduces object hallucination rates by approximately 19% compared to state-of-the-art baselines, while preserving general reasoning capabilities.

Keywords

Cite

@article{arxiv.2602.11824,
  title  = {Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models},
  author = {Jialin Wu and Wei Shi and Han Shen and Peigui Qi and Kunsheng Tang and Zhicong Huang and Binghao Wang and Zhou Yang},
  journal= {arXiv preprint arXiv:2602.11824},
  year   = {2026}
}

Comments

Accepted by ICML 2026

R2 v1 2026-07-01T10:33:28.182Z