English
Related papers

Related papers: ProstaTD: Bridging Surgical Triplet from Classific…

200 papers

Recently, there has been growing interest in developing learning-based methods to detect and utilize salient semi-global or global structures, such as junctions, lines, planes, cuboids, smooth surfaces, and all types of symmetries, for 3D…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Jia Zheng , Junfei Zhang , Jing Li , Rui Tang , Shenghua Gao , Zihan Zhou

We introduce a new large-scale data set of video URLs with densely-sampled object bounding box annotations called YouTube-BoundingBoxes (YT-BB). The data set consists of approximately 380,000 video segments about 19s long, automatically…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Esteban Real , Jonathon Shlens , Stefano Mazzocchi , Xin Pan , Vincent Vanhoucke

The scarcity of high-quality, logically annotated video datasets remains a primary bottleneck in advancing Multi-Modal Large Language Models (MLLMs) for the medical domain. Traditional manual annotation is prohibitively expensive and…

Artificial Intelligence · Computer Science 2025-12-02 Shenxi Liu , Kan Li , Mingyang Zhao , Yuhang Tian , Shoujun Zhou , Bin Li

3D object detection plays a crucial role in various applications such as autonomous vehicles, robotics and augmented reality. However, training 3D detectors requires a costly precise annotation, which is a hindrance to scaling annotation to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Saad Lahlali , Nicolas Granger , Hervé Le Borgne , Quoc-Cuong Pham

Video text spotting refers to localizing, recognizing, and tracking textual elements such as captions, logos, license plates, signs, and other forms of text within consecutive video frames. However, current datasets available for this task…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Haibin He , Jing Zhang , Mengyang Xu , Juhua Liu , Bo Du , Dacheng Tao

Minimally invasive image-guided surgery heavily relies on vision. Deep learning models for surgical video analysis could therefore support visual tasks such as assessing the critical view of safety (CVS) in laparoscopic cholecystectomy…

Image and Video Processing · Electrical Eng. & Systems 2021-09-21 Pietro Mascagni , Deepak Alapatt , Alain Garcia , Nariaki Okamoto , Armine Vardazaryan , Guido Costamagna , Bernard Dallemagne , Nicolas Padoy

Surgical tool detection is essential for analyzing and evaluating minimally invasive surgery videos. Current approaches are mostly based on supervised methods that require large, fully instance-level labels (i.e., bounding boxes). However,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Ryo Fujii , Ryo Hachiuma , Hideo Saito

Deep learning has achieved significant advancements in medical image segmentation. Currently, obtaining accurate segmentation outcomes is critically reliant on large-scale datasets with high-quality annotations. However, noisy annotations…

Image and Video Processing · Electrical Eng. & Systems 2026-01-08 Yuyang Fu , Xiuzhen Guo , Ji Shi

Teeth segmentation is an essential task in dental image analysis for accurate diagnosis and treatment planning. While supervised deep learning methods can be utilized for teeth segmentation, they often require extensive manual annotation of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Tomáš Kunzo , Viktor Kocur , Lukáš Gajdošech , Martin Madaras

Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (ZOD) provides rich multi-sensor data but only bounding-box labels, limiting its use for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Toomas Tahves , Mauro Bellone , Junyi Gu , Raivo Sell

Medical image analysis using deep learning is often challenged by limited labeled data and high annotation costs. Fine-tuning the entire network in label-limited scenarios can lead to overfitting and suboptimal performance. Recently, prompt…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Fan Bai , Ke Yan , Xiaoyu Bai , Xinyu Mao , Xiaoli Yin , Jingren Zhou , Yu Shi , Le Lu , Max Q. -H. Meng

Pixel-wise segmentation is one of the most data and annotation hungry tasks in our field. Providing representative and accurate annotations is often mission-critical especially for challenging medical applications. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Simon Reiß , Constantin Seibold , Alexander Freytag , Erik Rodner , Rainer Stiefelhagen

Existing medical imaging datasets for abdominal CT often lack three-dimensional annotations, multi-organ coverage, or precise lesion-to-organ associations, hindering robust representation learning and clinical applications. To address this…

Image and Video Processing · Electrical Eng. & Systems 2026-02-16 Mehran Advand , Zahra Dehghanian , Navid Faraji , Reza Barati , Seyed Amir Ahmad Safavi-Naini , Hamid R. Rabiee

Surgical scene segmentation is essential for anatomy and instrument localization which can be further used to assess tissue-instrument interactions during a surgical procedure. In 2017, the Challenge on Automatic Tool Annotation for…

Deep learning requires large amounts of data, and a well-defined pipeline for labeling and augmentation. Current solutions support numerous computer vision tasks with dedicated annotation types and formats, such as bounding boxes, polygons,…

Robotics · Computer Science 2023-12-01 G. Sharma , A. Angleraud , R. Pieters

Large labeled data sets are one of the essential basics of modern deep learning techniques. Therefore, there is an increasing need for tools that allow to label large amounts of data as intuitively as possible. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Dennis Stumpf , Stephan Krauß , Gerd Reis , Oliver Wasenmüller , Didier Stricker

In recent years, deep learning (DL) methods have become powerful tools for biomedical image segmentation. However, high annotation efforts and costs are commonly needed to acquire sufficient biomedical training data for DL models. To…

Computer Vision and Pattern Recognition · Computer Science 2018-06-05 Lin Yang , Yizhe Zhang , Zhuo Zhao , Hao Zheng , Peixian Liang , Michael T. C. Ying , Anil T. Ahuja , Danny Z. Chen

Video analysis has been moving towards more detailed interpretation (e.g. segmentation) with encouraging progresses. These tasks, however, increasingly rely on densely annotated training data both in space and time. Since such annotation is…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Yuzheng Xu , Yang Wu , Nur Sabrina binti Zuraimi , Shohei Nobuhara , Ko Nishino

Segmentation is one of the most important tasks in the medical imaging pipeline as it influences a number of image-based decisions. To be effective, fully supervised segmentation approaches require large amounts of manually annotated…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Tyler Ward , Aaron Moseley , Abdullah-Al-Zubaer Imran

The field of medical image segmentation is hindered by the scarcity of large, publicly available annotated datasets. Not all datasets are made public for privacy reasons, and creating annotations for a large dataset is time-consuming and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Iira Häkkinen , Iaroslav Melekhov , Erik Englesson , Hossein Azizpour , Juho Kannala