中文
相关论文

相关论文: OphNet: A Large-Scale Video Benchmark for Ophthalm…

200 篇论文

In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yujin Wang , Yutong Zheng , Wenxian Fan , Tianyi Wang , Hongqing Chu , Li Zhang , Bingzhao Gao , Daxin Tian , Jianqiang Wang , Hong Chen

Surgical triplet detection is a critical task in surgical video analysis. However, existing datasets like CholecT50 lack precise spatial bounding box annotations, rendering triplet classification at the image level insufficient for…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yiliang Chen , Zhixi Li , Cheng Xu , Alex Qinyang Liu , Ruize Cui , Xuemiao Xu , Jeremy Yuen-Chun Teoh , Shengfeng He , Jing Qin

Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain challenging due to the difficulty of aligning dense spatio-temporal video representations…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Kedi Sun , Chaohui Dang , Yue Feng , James Glasbey , Theodoros N. Arvanitis , Le Zhang

Mapping surgery is fundamental to developing operative guidelines and enabling autonomous robotic surgery. Recent advances in artificial intelligence (AI) have shown promise in mapping the behaviour of surgeons from videos, yet current…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dani Kiyasseh

Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events.…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Tongtong Yuan , Xuange Zhang , Kun Liu , Bo Liu , Chen Chen , Jian Jin , Zhenzhen Jiao

Autonomous robotic surgery has advanced significantly based on analysis of visual and temporal cues in surgical workflow, but relational cues from domain knowledge remain under investigation. Complex relations in surgical annotations can be…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Shang Zhao , Yanzhe Liu , Qiyuan Wang , Dai Sun , Rong Liu , S. Kevin Zhou

With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Yulin Fei , Yuhui Gao , Xingyuan Xian , Xiaojin Zhang , Tao Wu , Wei Chen

Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Yuxin He , An Li , Cheng Xue

This paper investigates the automatic monitoring of tool usage during a surgery, with potential applications in report generation, surgical training and real-time decision support. Two surgeries are considered: cataract surgery, the most…

计算机视觉与模式识别 · 计算机科学 2018-05-08 Hassan Al Hajj , Mathieu Lamard , Pierre-Henri Conze , Béatrice Cochener , Gwenolé Quellec

Optical flow estimation has been a long-lasting and fundamental problem in the computer vision community. However, despite the advances of optical flow estimation in perspective videos, the 360$^\circ$ videos counterpart remains in its…

计算机视觉与模式识别 · 计算机科学 2023-01-30 Bin Duan , Keshav Bhandari , Gaowen Liu , Yan Yan

We propose a new multilabel classifier, called LapTool-Net to detect the presence of surgical tools in each frame of a laparoscopic video. The novelty of LapTool-Net is the exploitation of the correlation among the usage of different tools…

计算机视觉与模式识别 · 计算机科学 2019-05-23 Babak Namazi , Ganesh Sankaranarayanan , Venkat Devarajan

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

Safety and efficiency are paramount in healthcare facilities where the lives of patients are at stake. Despite the adoption of robots to assist medical staff in challenging tasks such as complex surgeries, human expertise is still…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Rohit Mohan , José Arce , Sassan Mokhtar , Daniele Cattaneo , Abhinav Valada

In cataract surgery, the operation is performed with the help of a microscope. Since the microscope enables watching real-time surgery by up to two people only, a major part of surgical training is conducted using the recorded videos. To…

计算机视觉与模式识别 · 计算机科学 2021-09-28 Negin Ghamsarian , Mario Taschwer , Doris Putzgruber-Adamitsch , Stephanie Sarny , Klaus Schoeffmann

Endoscopic surgery is currently an important treatment method in the field of spinal surgery and avoiding damage to the spinal nerves through video guidance is a key challenge. This paper presents the first real-time segmentation method for…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Shaowu Peng , Pengcheng Zhao , Yongyu Ye , Junying Chen , Yunbing Chang , Xiaoqing Zheng

Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their…

Accurate tool tracking is essential for the success of computer-assisted intervention. Previous efforts often modeled tool trajectories rigidly, overlooking the dynamic nature of surgical procedures, especially tracking scenarios like…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Chinedu Innocent Nwoye , Nicolas Padoy

Automatic tool detection from surgical imagery has a multitude of useful applications, such as real-time computer assistance for the surgeon. Using the successful residual network architecture, a system that can distinguish 21 different…

计算机视觉与模式识别 · 计算机科学 2018-05-16 Jonas Prellberg , Oliver Kramer

The rapid progress of photorealistic synthesis techniques has reached at a critical point where the boundary between real and manipulated images starts to blur. Thus, benchmarking and advancing digital forgery analysis have become a…

计算机视觉与模式识别 · 计算机科学 2021-07-15 Yinan He , Bei Gan , Siyu Chen , Yichun Zhou , Guojun Yin , Luchuan Song , Lu Sheng , Jing Shao , Ziwei Liu

Humans possess the cognitive ability to comprehend scenes in a compositional manner. To empower AI systems with similar capabilities, object-centric learning aims to acquire representations of individual objects from visual scenes without…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Yinxuan Huang , Tonglin Chen , Zhimeng Shen , Jinghao Huang , Bin Li , Xiangyang Xue