中文

想象替代方案:基于语言引导的高分辨率3D反事实医学图像生成

图像与视频处理 2025-10-02 v2 计算与语言 计算机视觉与模式识别 机器学习

摘要

视觉语言模型在 various conditions 下生成2D图像方面展现了惊人的能力;然而,这些模型的成功很大程度上依赖于大量可获取的预训练基础模型。关键的是,针对3D的相当预训练模型并不存在,这显著限制了进展。结果,视觉语言模型在仅仅依靠自然语言条件下生成高分辨率3D反事实医学图像的潜力尚未被探索。弥补这一差距将 enables powerful 的临床和研究应用,例如个性化反事实解释、模拟疾病进展,以及通过在真实细节中可视化假设条件来增强医学培训。我们的工作采取一步方法,通过引入一个能够生成由 free-form language prompts 指导的合成患者的高分辨率3D反事实医学图像的框架。我们采用最新的3D扩散模型,并引入 Simple Diffusion 的改进,并 incorporates augmented conditioning 以 improve text alignment 和 image quality。鉴于本是首次演示的 language-guided native-3D diffusion model 应用于神经影像,其中忠实的三维建模是 essential 的。在两个神经影像 MRI 数据集上,我们的框架模拟了 Multiple Sclerosis 中的 varying counterfactual lesion loads 以及 Alzheimer's disease 中的认知状态,生成高质量图像,同时保持 subject fidelity。我们的结果为 prompt-driven disease progression analysis in 3D medical imaging 奠定了基础。项目链接 - https://lesupermomo.github.io/imagining-alternatives/。

关键词

引用

@article{arxiv.2509.05978,
  title  = {Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance},
  author = {Mohamed Mohamed and Brennan Nichyporuk and Douglas L. Arnold and Tal Arbel},
  journal= {arXiv preprint arXiv:2509.05978},
  year   = {2025}
}

备注

Accepted to the 2025 MICCAI ELAMI Workshop