中文

面向机器人手术任务的预测模型学习:从缝合世界模型入门

计算机视觉与模式识别 2025-03-18 v1

摘要

我们引入了专门的扩散基于生成模型,通过对标注后的腔腔动手术影像进行监督学习,捕捉细粒度机器人手术亚粘针动作的时空动力学。 proposed models form a foundation for data-driven world models capable of simulating the biomechanical interactions and procedural dynamics of surgical suturing with high temporal fidelity. Annotating a dataset of 2K\sim2K clips extracted from simulation videos, we categorize surgical actions into fine-grained sub-stitch classes including ideal and non-ideal executions of needle positioning, targeting, driving, and withdrawal. We fine-tune two state-of-the-art video diffusion models, LTX-Video and HunyuanVideo, to generate high-fidelity surgical action sequences at \ge768x512 resolution and \ge49 frames. For training our models, we explore both Low-Rank Adaptation (LoRA) and full-model fine-tuning approaches. Our experimental results demonstrate that these world models can effectively capture the dynamics of suturing, potentially enabling improved training simulators, surgical skill assessment tools, and autonomous surgical systems. The models also display the capability to differentiate between ideal and non-ideal technique execution, providing a foundation for building surgical training and evaluation systems. We release our models for testing and as a foundation for future research. Project Page: https://mkturkcan.github.io/suturingmodels/

关键词

引用

@article{arxiv.2503.12531,
  title  = {Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks},
  author = {Mehmet Kerem Turkcan and Mattia Ballo and Filippo Filicori and Zoran Kostic},
  journal= {arXiv preprint arXiv:2503.12531},
  year   = {2025}
}