中文
相关论文

相关论文: Open surgery tool classification and hand utilizat…

200 篇论文

Incorporating multiple camera views for detection alleviates the impact of occlusions in crowded scenes. In a multiview system, we need to answer two important questions when dealing with ambiguities that arise from occlusions. First, how…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Yunzhong Hou , Liang Zheng , Stephen Gould

We propose a novel multi-modal and multi-task architecture for simultaneous low level gesture and surgical task classification in Robot Assisted Surgery (RAS) videos.Our end-to-end architecture is based on the principles of a long…

计算机视觉与模式识别 · 计算机科学 2018-05-03 Duygu Sarikaya , Khurshid A. Guru , Jason J. Corso

3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera intrinsic ambiguity: the same image admits different 3D scenes under different cameras.…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Xueying Jiang , Wenhao Li , Quanhao Qian , Deli Zhao , Shijian Lu , Gongjie Zhang , Ran Xu

In this work, we explore whether it is possible to learn representations of endoscopic video frames to perform tasks such as identifying surgical tool presence without supervision. We use a maximum mean discrepancy (MMD) variational…

计算机视觉与模式识别 · 计算机科学 2020-08-31 David Z. Li , Masaru Ishii , Russell H. Taylor , Gregory D. Hager , Ayushi Sinha

Reliable depth estimation under real optical conditions remains a core challenge for camera vision in systems such as autonomous robotics and augmented reality. Despite recent progress in depth estimation and depth-of-field rendering,…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Nisarg K. Trivedi , Vinayak A. Belludi , Li-Yun Wang

In the developing countries, most of the Manual Material Handling (MMH) related tasks are labor-intensive. It is not possible for these countries to assess injuries during lifting heavy weight, as multi-camera motion capture, force plate…

医学物理 · 物理学 2021-04-01 Mahmudul Hasan , Miftahur Rahman

Precise calibration is the basis for the vision-guided robot system to achieve high-precision operations. Systems with multiple eyes (cameras) and multiple hands (robots) are particularly sensitive to calibration errors, such as…

机器人学 · 计算机科学 2023-05-05 Zishun Zhou , Liping Ma , Xilong Liu , Zhiqiang Cao , Junzhi Yu

During surgeries, there is a risk of medical gauzes being left inside patients' bodies, leading to "Gossypiboma" in patients and can cause serious complications in patients and also lead to legal problems for hospitals from malpractice…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Saraf Krish , Cai Yiyu , Huang Li Hui

Human-robot collaboration requires the establishment of methods to guarantee the safety of participating operators. A necessary part of this process is ensuring reliable human pose estimation. Established vision-based modalities encounter…

机器人学 · 计算机科学 2024-06-28 Michael Zechmair , Yannick Morel

Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Shiyu Li , Hannah Schieber , Kristoffer Waldow , Benjamin Busam , Julian Kreimeier , Daniel Roth

Camera calibration methods usually consist of capturing images of known calibration patterns and using the detected correspondences to optimize the parameters of the assumed camera model. A meaningful evaluation of these methods relies on…

计算机视觉与模式识别 · 计算机科学 2022-03-10 Tim Michels , Arne Petersen , Reinhard Koch

Multispectral imaging is very beneficial in diverse applications, like healthcare and agriculture, since it can capture absorption bands of molecules in different spectral areas. A promising approach for multispectral snapshot imaging are…

图像与视频处理 · 电气工程与系统科学 2024-11-08 Frank Sippel , Jürgen Seiler , André Kaup

Privacy preservation is a prerequisite for using video data in Operating Room (OR) research. Effective anonymization relies on the exhaustive localization of every individual; even a single missed detection necessitates extensive manual…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Keqi Chen , Vinkle Srivastav , Armine Vardazaryan , Cindy Rolland , Didier Mutter , Nicolas Padoy

The growing need for video surveillance in public spaces has created a demand for systems that can track individuals across multiple cameras feeds in real-time. While existing tracking systems have achieved impressive performance using deep…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Vipin Gautam , Shitala Prasad , Sharad Sinha

Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation system are crucial for ensuring high-quality, diverse, and…

机器人学 · 计算机科学 2024-07-09 Xuxin Cheng , Jialong Li , Shiqi Yang , Ge Yang , Xiaolong Wang

Accurate counting of surgical instruments in Operating Rooms (OR) is a critical prerequisite for ensuring patient safety during surgery. Despite recent progress of large visual-language models and agentic AI, accurately counting such…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Rishikesh Bhyri , Brian R Quaranto , Philip J Seger , Kaity Tung , Brendan Fox , Gene Yang , Steven D. Schwaitzberg , Junsong Yuan , Nan Xi , Peter C W Kim

For object detection in wide-area aerial imagery, post-processing is usually needed to reduce false detections. We propose a two-stage post-processing scheme which comprises an area-thresholding sieving process and a morphological closing…

计算机视觉与模式识别 · 计算机科学 2020-10-30 Xin Gao , Sundaresh Ram , Jeffrey J. Rodriguez

Anomaly detection in surveillance videos remains a challenging task due to the diversity of abnormal events, class imbalance, and scene-dependent visual clutter. To address these issues, we propose a robust deep learning framework that…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Mohammad Ali Etemadi Naeen , Hoda Mohammadzade , Saeed Bagheri Shouraki

This paper presents two novel approaches for people counting in crowded and open environments that combine the information gathered by multiple views. Multiple camera are used to expand the field of view as well as to mitigate the problem…

计算机视觉与模式识别 · 计算机科学 2017-05-09 Fabio Dittrich , Luiz E. S. de Oliveira , Alceu S. Britto , Alessandro L. Koerich

The UN-Habitat estimates that over one billion people live in slums around the world. However, state-of-the-art techniques to detect the location of slum areas employ high-resolution satellite imagery, which is costly to obtain and process.…

计算机视觉与模式识别 · 计算机科学 2021-06-23 Agatha C. H. de Mattos , Gavin McArdle , Michela Bertolotto