中文
相关论文

相关论文: SWoMo: Neuro-Symbolic World Model for Cataract Sur…

200 篇论文

A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI and simulation -- from data scarcity and rare event synthesis…

Realistic and interactive surgical simulation has the potential to facilitate crucial applications, such as medical professional training and autonomous surgical agent training. In the natural visual domain, world models have enabled…

图像与视频处理 · 电气工程与系统科学 2025-12-16 Saurabh Koju , Saurav Bastola , Prashant Shrestha , Sanskar Amgain , Yash Raj Shrestha , Rudra P. K. Poudel , Binod Bhattarai

We introduce specialized diffusion-based generative models that capture the spatiotemporal dynamics of fine-grained robotic surgical sub-stitch actions through supervised learning on annotated laparoscopic surgery footage. The proposed…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Mehmet Kerem Turkcan , Mattia Ballo , Filippo Filicori , Zoran Kostic

Surgical simulation is an increasingly important element of surgical education. Using simulation can be a means to address some of the significant challenges in developing surgical skills with limited time and resources. The photo-realistic…

计算机视觉与模式识别 · 计算机科学 2018-11-08 Imanol Luengo , Evangello Flouty , Petros Giataganas , Piyamate Wisanuvej , Jean Nehme , Danail Stoyanov

While neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal…

计算机视觉与模式识别 · 计算机科学 2021-10-04 Zhuowan Li , Elias Stengel-Eskin , Yixiao Zhang , Cihang Xie , Quan Tran , Benjamin Van Durme , Alan Yuille

This paper introduces a SSSUMO, semi-supervised deep learning approach for submovement decomposition that achieves state-of-the-art accuracy and speed. While submovement analysis offers valuable insights into motor control, existing methods…

人机交互 · 计算机科学 2025-07-14 Evgenii Rudakov , Jonathan Shock , Otto Lappi , Benjamin Ultan Cowley

Automation holds the potential to assist surgeons in robotic interventions, shifting their mental work load from visuomotor control to high level decision making. Reinforcement learning has shown promising results in learning complex…

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Execute" approach, which uses video prediction to infer actions…

机器人学 · 计算机科学 2026-05-12 Qiuxuan Feng , Jiale Yu , Jiaming Liu , Yueru Jia , Zhuangzhe Wu , Hao Chen , Zezhong Qian , Shuo Gu , Peng Jia , Siwei Ma , Shanghang Zhang

Real-time prediction of technical errors from cataract surgical videos can be highly beneficial, particularly for telementoring, which involves remote guidance and mentoring through digital platforms. However, the rarity of surgical errors…

图像与视频处理 · 电气工程与系统科学 2025-03-31 Maxime Faure , Pierre-Henri Conze , Béatrice Cochener , Anas-Alexis Benyoussef , Mathieu Lamard , Gwenolé Quellec

Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that generates high-fidelity, motion-controllable surgical videos…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Tong Chen , Shuya Yang , Junyi Wang , Long Bai , Hongliang Ren , Luping Zhou

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that…

机器人学 · 计算机科学 2025-09-18 Guanxing Lu , Baoxiong Jia , Puhao Li , Yixin Chen , Ziwei Wang , Yansong Tang , Siyuan Huang

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Yufu Wang , Evonne Ng , Soyong Shin , Rawal Khirodkar , Yuan Dong , Zhaoen Su , Jinhyung Park , Kris Kitani , Alexander Richard , Fabian Prada , Michael Zollhofer

This paper tackles the challenge of automatically performing realistic surgical simulations from readily available surgical videos. Recent efforts have successfully integrated physically grounded dynamics within 3D Gaussians to perform…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Kailing Wang , Chen Yang , Keyang Zhao , Xiaokang Yang , Wei Shen

World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision-making. Yet, despite rapid progress in industry-scale interactive video generation, the broader research…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Siqiao Huang , Partha Kaushik , Michael Chen , Hengkai Pan , Kaiwen Geng , Omar Chehab , Fernando Moreno-Pino , Max Simchowitz

Recent advances in robotic learning in simulation have shown impressive results in accelerating learning complex manipulation skills. However, the sim-to-real gap, caused by discrepancies between simulation and reality, poses significant…

机器人学 · 计算机科学 2025-03-25 Jacinto Colan , Keisuke Sugita , Ana Davila , Yutaro Yamada , Yasuhisa Hasegawa

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Sen Wang , Jingyi Tian , Le Wang , Zhimin Liao , Jiayi Li , Huaiyi Dong , Kun Xia , Sanping Zhou , Wei Tang , Hua Gang

Dynamic human rendering from video sequences has achieved remarkable progress by formulating the rendering as a mapping from static poses to human images. However, existing methods focus on the human appearance reconstruction of every…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Tao Hu , Fangzhou Hong , Ziwei Liu

Prior work for articulated 3D shape reconstruction often relies on specialized sensors (e.g., synchronized multi-camera systems), or pre-built 3D deformable models (e.g., SMAL or SMPL). Such methods are not able to scale to diverse sets of…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Gengshan Yang , Minh Vo , Natalia Neverova , Deva Ramanan , Andrea Vedaldi , Hanbyul Joo

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yujiang Pu , Zhanbo Huang , Vishnu Boddeti , Yu Kong

An embodied system must not only model the patterns of the external world but also understand its own motion dynamics. A motion dynamic model is essential for efficient skill acquisition and effective planning. In this work, we introduce…

机器学习 · 计算机科学 2025-04-10 Chenjie Hao , Weyl Lu , Yifan Xu , Yubei Chen
‹ 上一页 1 2 3 10 下一页 ›