中文

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

计算机视觉与模式识别 2026-05-29 v1

摘要

多模态检索增强生成(MMRAG)已成为增强多模态大型语言模型在知识密集型问答中表现的强大范式,通过整合外部视觉、文本和结构知识。然而,现有 MMRAG 框架存在关键局限,包括噪声和无关检索、跨模态语义错位、缺乏自适应推理以及局部与全局语境之间的生成不连贯。我们引入 CogniVerse,一种新型 MMRAG 框架,通过认知启发的、数学严谨的方法来解决这些挑战。drawing from human-like reasoning, CogniVerse integrates three synergistic components: (1) a Cognitive Reflection Module that dynamically assesses retrieval necessity and filters relevant multi-modal content, reducing noise and computational overhead; (2) a Multi-modal Retrieval Module that aligns embeddings in a Riemannian manifold using information geometry and refines knowledge graphs via spectral graph theory, ensuring precise and coherent retrieval; (3) a Hierarchical Generation Module that employs an optimal transport-based loss to balance token-level accuracy and global semantic coherence. Extensive experiments demonstrate that CogniVerse significantly outperforms state-of-the-art systems in both accuracy and coherence, while reducing retrieval latency.

关键词

引用

@article{arxiv.2605.29602,
  title  = {CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning},
  author = {Xiang Fang and Wanlong Fang and Changshuo Wang},
  journal= {arXiv preprint arXiv:2605.29602},
  year   = {2026}
}

备注

Accepted in CVPR 2026