中文
相关论文

相关论文: MM-OR: A Large Multimodal Operating Room Dataset f…

200 篇论文

Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their development is limited by the lack of realistic training and test environments: Real data is…

Mapping and scene representation are fundamental to reliable planning and navigation in mobile robots. While purely geometric maps using voxel grids allow for general navigation, obtaining up-to-date spatial and semantically rich…

机器人学 · 计算机科学 2025-03-12 Tim Steinke , Martin Büchner , Niclas Vödisch , Abhinav Valada

A high-fidelity digital simulation environment is crucial for accurately replicating physical operational processes. However, inconsistencies between simulation and physical environments result in low confidence in simulation outcomes,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Songjie Han , Yinhua Liu , Yanzheng Li , Hua Chen , Dongmei Yang

Hand-surface interactions between clinicians, patients, and medical equipment play a central role in pathogen transmission during medical procedures. However, these interactions remain largely unobserved, as current infection-prevention…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Sophokles Ktistakis , Rui Wang , Bastian Grande , Hugo Sax

Our goal is to develop stable, accurate, and robust semantic scene understanding methods for wide-area scene perception and understanding, especially in challenging outdoor environments. To achieve this, we are exploring and evaluating a…

计算机视觉与模式识别 · 计算机科学 2021-04-01 Jiesi Hu , Ganning Zhao , Suya You , C. C. Jay Kuo

Every day, countless surgeries are performed worldwide, each within the distinct settings of operating rooms (ORs) that vary not only in their setups but also in the personnel, tools, and equipment used. This inherent diversity poses a…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Ege Özsoy , Chantal Pellegrini , Matthias Keicher , Nassir Navab

In recent years, 3D scene graphs have emerged as a powerful world representation, offering both geometric accuracy and semantic richness. Combining 3D scene graphs with large language models enables robots to reason, plan, and navigate in…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Abdelrhman Werby , Dennis Rotondi , Fabio Scaparro , Kai O. Arras

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

图像与视频处理 · 电气工程与系统科学 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

Automated identification of surgical safety risks is critical for improving patient outcomes; however, Multimodal Large Language Models (MLLMs) frequently suffer from Visual-Semantic Knowledge Conflicts (VS-KC), a phenomenon where models…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Weiyi Zhao , Xiaoyu Tan , Liang Liu , Sijia Li , Youwei Song , Xihe Qiu

Purpose: Recent advances in Surgical Data Science (SDS) have contributed to an increase in video recordings from hospital environments. While methods such as surgical workflow recognition show potential in increasing the quality of patient…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Lennart Bastian , Tony Danjun Wang , Tobias Czempiel , Benjamin Busam , Nassir Navab

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR interpretation, enhancing diagnostic accuracy and efficiency.…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Qingqiu Li , Zihang Cui , Seongsu Bae , Jilan Xu , Runtian Yuan , Yuejie Zhang , Rui Feng , Quanli Shen , Xiaobo Zhang , Junjun He , Shujun Wang

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment. Ensuring its safety on public roads is one of the main duties that requires an accurate perception of the environment. To achieve…

Purpose: The integration of multimodal imaging into operating rooms paves the way for comprehensive surgical scene understanding. In ophthalmic surgery, by now, two complementary imaging modalities are available: operating microscope (OPMI)…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Nikolo Rohrmoser , Ghazal Ghazaei , Michael Sommersperger , Nassir Navab

The advent of generalist Large Language Models (LLMs) and Large Vision Models (VLMs) have streamlined the construction of semantically enriched maps that can enable robots to ground high-level reasoning and planning into their…

机器人学 · 计算机科学 2024-11-06 Emilio Olivastri , Jonathan Francis , Alberto Pretto , Niko Sünderhauf , Krishan Rana

Deep neural networks are prone to learning spurious correlations, exploiting dataset-specific artifacts rather than meaningful features for prediction. In surgical operating rooms (OR), these manifest through the standardization of smocks…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Tony Danjun Wang , Tobias Czempiel , Nassir Navab , Lennart Bastian

Understanding symptom-image associations is crucial for clinical reasoning. However, existing medical multimodal models often rely on simple one-to-one hard labeling, oversimplifying clinical reality where symptoms relate to multiple…

计算机视觉与模式识别 · 计算机科学 2025-11-11 You-Kyoung Na , Yeong-Jun Cho

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However, most scene graph mapping methods rely on depth cameras or LiDAR sensors. In this work, we…

机器人学 · 计算机科学 2026-05-14 Christina Kassab , Hyeonjae Gil , Matías Mattamala , Ayoung Kim , Maurice Fallon