中文
相关论文

相关论文: How Far Are Surgeons from Surgical World Models? A…

200 篇论文

Surgical video understanding is essential for computer-assisted interventions, yet existing surgical foundation models remain constrained by limited data scale, procedural diversity, and inconsistent evaluation, often lacking a reproducible…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Sicheng Lu , Zikai Xiao , Jianhui Wei , Danyu Sun , Qi Lu , Keli Hu , Yang Feng , Jian Wu , Zongxin Yang , Zuozhu Liu

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular reflections, and…

Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that generates high-fidelity, motion-controllable surgical videos…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Tong Chen , Shuya Yang , Junyi Wang , Long Bai , Hongliang Ren , Luping Zhou

Surgical video understanding is pivotal for enabling automated intraoperative decision-making, skill assessment, and postoperative quality improvement. However, progress in developing surgical video foundation models (FMs) remains hindered…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Jianhui Wei , Zikai Xiao , Danyu Sun , Luqi Gong , Zongxin Yang , Zuozhu Liu , Jian Wu

Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Cheng Yuan , Jian Jiang , Kunyi Yang , Lv Wu , Rui Wang , Zi Meng , Haonan Ping , Ziyu Xu , Yifan Zhou , Wanli Song , Hesheng Wang , Yueming Jin , Qi Dou , Yutong Ban

Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these…

其他定量生物学 · 定量生物学 2026-01-21 Yaoqian Li , Xikai Yang , Dunyuan Xu , Yang Yu , Litao Zhao , Xiaowei Hu , Jinpeng Li , Pheng-Ann Heng

Realistic and interactive surgical simulation has the potential to facilitate crucial applications, such as medical professional training and autonomous surgical agent training. In the natural visual domain, world models have enabled…

图像与视频处理 · 电气工程与系统科学 2025-12-16 Saurabh Koju , Saurav Bastola , Prashant Shrestha , Sanskar Amgain , Yash Raj Shrestha , Rudra P. K. Poudel , Binod Bhattarai

The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. This transformation emerged from simple primitives: large,…

Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-shot generalization across domains. Motivated by this trend, we investigate whether…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Yuxiang Lai , Jike Zhong , Ming Li , Yuheng Li , Xiaofeng Yang

Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Alejandra Perez , Anita Rau , Lee White , Busisiwe Mlambo , Chinedu Nwoye , Muhammad Abdullah Jamal , Omid Mohareri

This paper tackles the challenge of automatically performing realistic surgical simulations from readily available surgical videos. Recent efforts have successfully integrated physically grounded dynamics within 3D Gaussians to perform…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Kailing Wang , Chen Yang , Keyang Zhao , Xiaokang Yang , Wei Shen

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Lennart Maack , Alexander Schlaefer

Videos are prominent learning materials to prepare surgical trainees before they enter the operating room (OR). In this work, we explore techniques to enrich the video-based surgery learning experience. We propose Surgment, a system that…

人机交互 · 计算机科学 2024-06-27 Jingying Wang , Haoran Tang , Taylor Kantor , Tandis Soltani , Vitaliy Popov , Xu Wang

A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI and simulation -- from data scarcity and rare event synthesis…

Controllable medical video generation has achieved remarkable progress, but it still lacks interpretability, which requires the alignment of generated contents with physical priors and faithful clinical manifestations. To push the…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Junhu Fu , Ke Chen , Weidong Guo , Shuyu Liang , Jie Xu , Chen Ma , Kehao Wang , Shengli Lin , Zeju Li , Yuanyuan Wang , Yi Guo , Shuo Li

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Zechen Bai , Hai Ci , Mike Zheng Shou

Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their…

Video generation models have advanced rapidly and are beginning to show a strong understanding of physical dynamics. In this paper, we investigate how far an advanced video generation model such as Veo-3 can support generalizable robotic…

机器人学 · 计算机科学 2026-04-07 Zhongru Zhang , Chenghan Yang , Qingzhou Lu , Yanjiang Guo , Jianke Zhang , Yucheng Hu , Jianyu Chen

Data scarcity remains a fundamental barrier to achieving fully autonomous surgical robots. While large scale vision language action (VLA) models have shown impressive generalization in household and industrial manipulation by leveraging…

Modeling disease progression is crucial for improving the quality and efficacy of clinical diagnosis and prognosis, but it is often hindered by a lack of longitudinal medical image monitoring for individual patients. To address this…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Xu Cao , Kaizhao Liang , Kuei-Da Liao , Tianren Gao , Wenqian Ye , Jintai Chen , Zhiguang Ding , Jianguo Cao , James M. Rehg , Jimeng Sun
‹ 上一页 1 2 3 10 下一页 ›