English
Related papers

Related papers: ProstaTD: Bridging Surgical Triplet from Classific…

200 papers

Advances in surgical video analysis are transforming operating rooms into intelligent, data-driven environments. Computer-assisted systems support full surgical workflow, from preoperative planning to intraoperative guidance and…

Image and Video Processing · Electrical Eng. & Systems 2025-09-22 Sahar Nasirihaghighi

Background: Deep learning has presented great potential in accurate MR image segmentation when enough labeled data are provided for network optimization. However, manually annotating 3D MR images is tedious and time-consuming, requiring…

Image and Video Processing · Electrical Eng. & Systems 2023-12-19 Yousuf Babiker M. Osman , Cheng Li , Weijian Huang , Shanshan Wang

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve safety within the…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Gaurav Yengera , Didier Mutter , Jacques Marescaux , Nicolas Padoy

An accurate prostate delineation and volume characterization can support the clinical assessment of prostate cancer. A large amount of automatic prostate segmentation tools consider exclusively the axial MRI direction in spite of the…

Image and Video Processing · Electrical Eng. & Systems 2023-09-18 Tim Nikolass Lindeijer , Tord Martin Ytredal , Trygve Eftestøl , Tobias Nordström , Fredrik Jäderling , Martin Eklund , Alvaro Fernandez-Quilez

Automatic segmentation of medical images is a key step for diagnostic and interventional tasks. However, achieving this requires large amounts of annotated volumes, which can be tedious and time-consuming task for expert annotators. In this…

With only bounding-box annotations in the spatial domain, existing video scene text detection (VSTD) benchmarks lack temporal relation of text instances among video frames, which hinders the development of video text-related applications.…

Computer Vision and Pattern Recognition · Computer Science 2020-11-20 Yuanqiang Cai , Chang Liu , Weiqiang Wang , Qixiang Ye

Annotating a large-scale in-the-wild person re-identification dataset especially of marathon runners is a challenging task. The variations in the scenarios such as camera viewpoints, resolution, occlusion, and illumination make the problem…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Pranjal Singh Rajput , Yeshwanth Napolean , Jan van Gemert

We introduce the largest abdominal CT dataset (termed AbdomenAtlas) of 20,460 three-dimensional CT volumes sourced from 112 hospitals across diverse populations, geographies, and facilities. AbdomenAtlas provides 673K high-quality masks of…

Indexing endoscopic surgical videos is vital in surgical data science, forming the basis for systematic retrospective analysis and clinical performance evaluation. Despite its significance, current video analytics rely on manual indexing, a…

Understanding a surgical scene is crucial for computer-assisted surgery systems to provide any intelligent assistance functionality. One way of achieving this scene understanding is via scene segmentation, where every pixel of a frame is…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Alexander C. Jenke , Sebastian Bodenstedt , Fiona R. Kolbinger , Marius Distler , Jürgen Weitz , Stefanie Speidel

Current state-of-the-art (SOTA) 3D object detection methods often require a large amount of 3D bounding box annotations for training. However, collecting such large-scale densely-supervised datasets is notoriously costly. To reduce the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Chenqiang Gao , Chuandong Liu , Jun Shu , Fangcen Liu , Jiang Liu , Luyu Yang , Xinbo Gao , Deyu Meng

Automatic medical image segmentation plays a critical role in scientific research and medical care. Existing high-performance deep learning methods typically rely on large training datasets with high-quality manual annotations, which are…

Image and Video Processing · Electrical Eng. & Systems 2021-11-17 Shanshan Wang , Cheng Li , Rongpin Wang , Zaiyi Liu , Meiyun Wang , Hongna Tan , Yaping Wu , Xinfeng Liu , Hui Sun , Rui Yang , Xin Liu , Jie Chen , Huihui Zhou , Ismail Ben Ayed , Hairong Zheng

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Yunfei Xie , Ce Zhou , Lang Gao , Juncheng Wu , Xianhang Li , Hong-Yu Zhou , Sheng Liu , Lei Xing , James Zou , Cihang Xie , Yuyin Zhou

Annotating medical images, particularly for organ segmentation, is laborious and time-consuming. For example, annotating an abdominal organ requires an estimated rate of 30-60 minutes per CT volume based on the expertise of an annotator and…

Image and Video Processing · Electrical Eng. & Systems 2025-07-09 Chongyu Qu , Tiezheng Zhang , Hualin Qiao , Jie Liu , Yucheng Tang , Alan Yuille , Zongwei Zhou

3D multi-object detection and tracking are crucial for traffic scene understanding. However, the community pays less attention to these areas due to the lack of a standardized benchmark dataset to advance the field. Moreover, existing…

Computer Vision and Pattern Recognition · Computer Science 2019-03-07 Abhishek Patil , Srikanth Malla , Haiming Gang , Yi-Ting Chen

Accurate tool tracking is essential for the success of computer-assisted intervention. Previous efforts often modeled tool trajectories rigidly, overlooking the dynamic nature of surgical procedures, especially tracking scenarios like…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Chinedu Innocent Nwoye , Nicolas Padoy

Surgical phase recognition is a key task in computer-assisted surgery, aiming to automatically identify and categorize the different phases within a surgical procedure. Despite substantial advancements, most current approaches rely on fully…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Or Rubin , Shlomi Laufer

Annotating object ground truth in videos is vital for several downstream tasks in robot perception and machine learning, such as for evaluating the performance of an object tracker or training an image-based object detector. The accuracy of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Eric Price , Aamir Ahmad

Prostate cancer is the most prevalent cancer among men in Western countries, with 1.1 million new diagnoses every year. The gold standard for the diagnosis of prostate cancer is a pathologists' evaluation of prostate tissue. To potentially…

Image and Video Processing · Electrical Eng. & Systems 2020-10-23 Hans Pinckaers , Wouter Bulten , Jeroen van der Laak , Geert Litjens