English

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

Computer Vision and Pattern Recognition 2026-05-13 v5

Abstract

Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing representation informativeness and reducing inference latency. By analyzing token-level gradient dynamics during latent training, we reveal two critical observations: (1) visual tokens exhibit significantly smaller gradient norms than their textual counterparts due to inherent language bias, resulting in systematic visual under-optimization; and (2) semantically simple tokens converge rapidly, whereas complex tokens exhibit persistent gradient instability constrained by fixed architectural depths. To address these limitations, we propose a visual replay module and routing depth scaling to collaboratively enhance visual perception and refine complicated latents for deeper contextual reasoning. The former module leverages causal self-attention to estimate token saliency, reinforcing fine-grained grounding through spatially-coherent constraints. Complementarily, the latter mechanism adaptively allocates additional reasoning steps to complex tokens, enabling deeper contextual refinement. Guided by a curriculum strategy that progressively internalizes explicit CoT into compact latent representations, our framework achieves state-of-the-art performance across diverse benchmarks while delivering substantial inference speedups over explicit CoT baselines.

Keywords

Cite

@article{arxiv.2604.10500,
  title  = {Visual Enhanced Depth Scaling for Multimodal Latent Reasoning},
  author = {Yudong Han and Yong Wang and Zaiquan Yang and Zhen Qu and Liyuan Pan and Xiangxiang Chu},
  journal= {arXiv preprint arXiv:2604.10500},
  year   = {2026}
}

Comments

11 pages, 6 figures

R2 v1 2026-07-01T12:04:49.065Z