中文

HAVIR:基于 CLIP 指导的层次化视觉到图像重建

计算机视觉与模式识别 2025-07-08 v2 人工智能

摘要

从脑活动重建视觉信息的任务桥接了神经科学与计算机视觉之间。尽管已取得一定进展用于解码 fMRI 以生成图像,但在准确恢复高度复杂视觉刺激方面仍面临挑战。这一困难源于其元素密度和多样性、复杂的空间结构以及多层次语义信息。为此,我们提出 HAVIR,包含两个适配器:(1)AutoKL 适配器将 fMRI 体素转化为潜在扩散先验,捕获拓扑结构;(2)CLIP 适配器将体素转换为 CLIP 文本和图像嵌入,包含语义信息。这些互补的表示通过 Versatile Diffusion 融合以生成最终重建图像。为从复杂情境中提取最本质的语义信息,CLIP 适配器使用描述视觉刺激的文本说明及其对应的语义图像进行训练。实验结果表明,HAVIR 在复杂情境下有效重建视觉刺激的结构特征和语义信息,超越现有模型。

关键词

引用

@article{arxiv.2506.06035,
  title  = {HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion},
  author = {Shiyi Zhang and Dong Liang and Hairong Zheng and Yihang Zhou},
  journal= {arXiv preprint arXiv:2506.06035},
  year   = {2025}
}

备注

We have decided to withdraw this paper because the baseline methods used for comparison are outdated and do not reflect the current state-of-the-art. This significantly affects the validity of our performance claims and conclusions. We plan to conduct a more comprehensive evaluation and submit a revised version in the future