中文
相关论文

相关论文: SCOPE: Speech-guided COllaborative PErception Fram…

200 篇论文

Surgical workflow anticipation can give predictions on what steps to conduct or what instruments to use next, which is an essential part of the computer-assisted intervention system for surgery, e.g. workflow reasoning in robotic surgery.…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Xiatian Zhang , Noura Al Moubayed , Hubert P. H. Shum

The accurate reconstruction of surgical scenes from surgical videos is critical for various applications, including intraoperative navigation and image-guided robotic surgery automation. However, previous approaches, mainly relying on depth…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Ange Lou , Yamin Li , Xing Yao , Yike Zhang , Jack Noble

This paper presents a collaborative implicit neural simultaneous localization and mapping (SLAM) system with RGB-D image sequences, which consists of complete front-end and back-end modules including odometry, loop detection, sub-map…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Jiarui Hu , Mao Mao , Hujun Bao , Guofeng Zhang , Zhaopeng Cui

Open-world point cloud semantic segmentation (OW-Seg) aims to predict point labels of both base and novel classes in real-world scenarios. However, existing methods rely on resource-intensive offline incremental learning or densely…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Peng Zhang , Songru Yang , Jinsheng Sun , Weiqing Li , Zhiyong Su

Holistic scene understanding poses a fundamental contribution to the autonomous operation of a robotic agent in its environment. Key ingredients include a well-defined representation of the surroundings to capture its spatial structure as…

机器人学 · 计算机科学 2024-05-24 Niclas Vödisch

Object manipulation requires accurate object pose estimation. In open environments, robots encounter unknown objects, which requires semantic understanding in order to generalize both to known categories and beyond. To resolve this…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Peter Hönig , Stefan Thalhammer , Jean-Baptiste Weibel , Matthias Hirschmanner , Markus Vincze

Unsupervised image segmentation is a critical task in computer vision. It enables dense scene understanding without human annotations, which is especially valuable in domains where labelled data is scarce. However, existing methods often…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Boujemaa Guermazi , Riadh Ksantini , Naimul Khan

Surgical image segmentation is highly challenging, primarily due to scarcity of annotated data. Generalist prompted segmentation models like the Segment-Anything Model (SAM) can help tackle this task, but because they require image-specific…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Aditya Murali , Farahdiba Zarin , Adrien Meyer , Pietro Mascagni , Didier Mutter , Nicolas Padoy

Current orthopedic robotic systems largely focus on navigation, aiding surgeons in positioning a guiding tube but still requiring manual drilling and screw placement. The automation of this task not only demands high precision and safety…

机器人学 · 计算机科学 2025-02-04 Chen Chen , Qikai Zou , Yuhang Song , Mingrui Yu , Senqiang Zhu , Shiji Song , Xiang Li

Medical image segmentation is vital for clinical diagnosis and quantitative analysis, yet remains challenging due to the heterogeneity of imaging modalities and the high cost of pixel-level annotations. Although general interactive…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Yujie Lu , Jingwen Li , Sibo Ju , Yanzhou Su , he yao , Yisong Liu , Min Zhu , Junlong Cheng

Zero-shot referring expression comprehension (REC) aims to locate target objects in images given natural language queries without relying on task-specific training data, demanding strong visual understanding capabilities. Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yike Wu , Necva Bolucu , Stephen Wan , Dadong Wang , Jiahao Xia , Jian Zhang

Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Kun Yuan , Vinkle Srivastav , Tong Yu , Joel L. Lavanchy , Jacques Marescaux , Pietro Mascagni , Nassir Navab , Nicolas Padoy

Autonomous exploration in unknown environments is key for mobile robots, helping them perceive, map, and make decisions in complex areas. However, current methods often rely on frequent global optimization, suffering from high computational…

机器人学 · 计算机科学 2026-02-27 Kai Li , Shengtao Zheng , Linkun Xiu , Yuze Sheng , Xiao-Ping Zhang , Dongyue Huang , Xinlei Chen

This work introduces the first framework for reconstructing surgical dialogue from unstructured real-world recordings, which is crucial for characterizing teaching tasks. In surgical training, the formative verbal feedback that trainers…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Firdavs Nasriddinov , Rafal Kocielnik , Arushi Gupta , Cherine Yang , Elyssa Wong , Anima Anandkumar , Andrew Hung

Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as surgery proceeds. In the operating room, critical decisions depend on subtle,…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Jingyi He , Yue Zhou , Long Bai , Kun Yuan , Nassir Navab , Yuan Bi

Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Cheng Yuan , Jian Jiang , Kunyi Yang , Lv Wu , Rui Wang , Zi Meng , Haonan Ping , Ziyu Xu , Yifan Zhou , Wanli Song , Hesheng Wang , Yueming Jin , Qi Dou , Yutong Ban

Autonomous driving systems remain brittle in rare, ambiguous, and out-of-distribution scenarios, where human driver succeed through contextual reasoning. Shared autonomy has emerged as a promising approach to mitigate such failures by…

机器人学 · 计算机科学 2025-11-07 Phat Nguyen , Erfan Aasi , Shiva Sreeram , Guy Rosman , Andrew Silva , Sertac Karaman , Daniela Rus

Surgical scene understanding is crucial for computer-assisted intervention systems, requiring visual comprehension of surgical scenes that involves diverse elements such as surgical tools, anatomical structures, and their interactions. To…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Jongmin Shin , Enki Cho , Ka Young Kim , Jung Yong Kim , Seong Tae Kim , Namkee Oh

The Segment Anything Model (SAM) and CLIP are remarkable vision foundation models (VFMs). SAM, a prompt driven segmentation model, excels in segmentation tasks across diverse domains, while CLIP is renowned for its zero shot recognition…

Most cognitive architectures rely on discrete representation, both in space (e.g., objects) and in time (e.g., events). However, a robot interaction with the world is inherently continuous, both in space and in time. The segmentation of the…

机器人学 · 计算机科学 2016-11-25 Bruno Nery , Rodrigo Ventura