中文

ColonCrafter:基于扩散先验的结肠镜视频深度估计模型

计算机视觉与模式识别 2025-09-18 v1 人工智能 机器学习

摘要

结肠镜三维(3D)场景理解 presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle with temporal consistency across video sequences, limiting their applicability for 3D reconstruction. We present ColonCrafter, a diffusion-based depth estimation model that generates temporally consistent depth maps from monocular colonoscopy videos. Our approach learns robust geometric priors from synthetic colonoscopy sequences to generate temporally consistent depth maps. We also introduce a style transfer technique that preserves geometric structure while adapting real clinical videos to match our synthetic training domain. ColonCrafter achieves state-of-the-art zero-shot performance on the C3VD dataset, outperforming both general-purpose and endoscopy-specific approaches. Although full trajectory 3D reconstruction remains a challenge, we demonstrate clinically relevant applications of ColonCrafter, including 3D point cloud generation and surface coverage assessment.

关键词

引用

@article{arxiv.2509.13525,
  title  = {ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors},
  author = {Romain Hardy and Tyler Berzin and Pranav Rajpurkar},
  journal= {arXiv preprint arXiv:2509.13525},
  year   = {2025}
}

备注

12 pages, 8 figures