English
Related papers

Related papers: How many samples to label for an application given…

200 papers

Recent success of large-scale pre-trained language models crucially hinge on fine-tuning them on large amounts of labeled data for the downstream task, that are typically expensive to acquire. In this work, we study self-training as one of…

Computation and Language · Computer Science 2020-06-30 Subhabrata Mukherjee , Ahmed Hassan Awadallah

Chest radiography (CXR) plays a crucial role in the diagnosis of various diseases. However, the inherent class imbalance in the distribution of clinical findings presents a significant challenge for current self-supervised deep learning…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Rajesh Madhipati , Sheethal Bhat , Lukas Buess , Andreas Maier

Chest X-ray (CXR) imaging is widely used for screening and diagnosing pulmonary abnormalities, yet automated interpretation remains challenging due to weak disease signals, dataset bias, and limited spatial supervision. Foundation models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Brayden Miao , Zain Rehman , Xin Miao , Siming Liu , Jianjie Wang

Machine learning models for radiology benefit from large-scale data sets with high quality labels for abnormalities. We curated and analyzed a chest computed tomography (CT) data set of 36,316 volumes from 19,993 unique patients. This is…

Image and Video Processing · Electrical Eng. & Systems 2020-10-14 Rachel Lea Draelos , David Dov , Maciej A. Mazurowski , Joseph Y. Lo , Ricardo Henao , Geoffrey D. Rubin , Lawrence Carin

Labelling data is a major practical bottleneck in training and testing classifiers. Given a collection of unlabelled data points, we address how to select which subset to label to best estimate test metrics such as accuracy, $F_1$ score or…

Machine Learning · Computer Science 2021-09-27 Emine Yilmaz , Peter Hayes , Raza Habib , Jordan Burgess , David Barber

Diagnostic imaging often requires the simultaneous identification of a multitude of findings of varied size and appearance. Beyond global indication of said findings, the prediction and display of localization information improves trust in…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Li Yao , Jordan Prosky , Eric Poblenz , Ben Covington , Kevin Lyman

A key challenge for the widespread application of learning-based models for robotic perception is to significantly reduce the required amount of annotated training data while achieving accurate predictions. This is essential not only to…

Robotics · Computer Science 2024-12-04 Niclas Vödisch , Kürsat Petek , Markus Käppeler , Abhinav Valada , Wolfram Burgard

Breast cancer prediction models for mammography assume that annotations are available for individual images or regions of interest (ROIs), and that there is a fixed number of images per patient. These assumptions do not hold in real…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Shreyasi Pathak , Jörg Schlötterer , Jeroen Geerdink , Jeroen Veltman , Maurice van Keulen , Nicola Strisciuglio , Christin Seifert

When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distributions subsampled from ChaosNLI, a…

Computation and Language · Computer Science 2026-05-29 Guneet Kohli

The current accessibility to large medical datasets for training convolutional neural networks is tremendously high. The associated dataset labels are always considered to be the real "ground truth". However, the labeling procedures often…

Computer Vision and Pattern Recognition · Computer Science 2019-12-13 Sebastian Guendel , Andreas Maier

Pretrained language models have improved zero-shot text classification by allowing the transfer of semantic knowledge from the training data in order to classify among specific label sets in downstream tasks. We propose a simple way to…

Computation and Language · Computer Science 2023-10-24 Lingyu Gao , Debanjan Ghosh , Kevin Gimpel

The interpretation of medical images is a challenging task, often complicated by the presence of artifacts, occlusions, limited contrast and more. Most notable is the case of chest radiography, where there is a high inter-rater variability…

Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on closed-set classes from a single institution, failing to…

Although deep learning can provide promising results in medical image analysis, the lack of very large annotated datasets confines its full potential. Furthermore, limited positive samples also create unbalanced datasets which limit the…

Computer Vision and Pattern Recognition · Computer Science 2018-05-09 Ken C. L. Wong , Alexandros Karargyris , Tanveer Syeda-Mahmood , Mehdi Moradi

Computed tomography (CT) is a key imaging modality for diagnosis, yet its clinical utility is marred by high radiation exposure and long turnaround times, restricting its use for larger-scale screening. Although chest radiography (CXR) is…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jianzhong You , Yuan Gao , Sangwook Kim , Chris Mcintosh

Chest X-Ray imaging is one of the most common radiological tools for detection of various pathologies related to the chest area and lung function. In a clinical setting, automated assessment of chest radiographs has the potential of…

Machine Learning · Computer Science 2022-10-31 David Biesner , Helen Schneider , Benjamin Wulff , Ulrike Attenberger , Rafet Sifa

In this work, we apply self-supervised learning with instance differentiation to learn a robust, multi-purpose representation for image analysis of resolved extragalactic continuum images. We train a multi-use model which compresses our…

Instrumentation and Methods for Astrophysics · Physics 2023-10-20 Inigo V. Slijepcevic , Anna M. M. Scaife , Mike Walmsley , Micah Bowles , O. Ivy Wong , Stanislav S. Shabala , Sarah V. White

Many real-world image recognition problems, such as diagnostic medical imaging exams, are "long-tailed" $\unicode{x2013}$ there are a few common findings followed by many more relatively rare conditions. In chest radiography, diagnosis is…

It is often infeasible or impossible to obtain ground truth labels for medical data. To circumvent this, one may build rule-based or other expert-knowledge driven labelers to ingest data and yield silver labels absent any ground-truth…

Machine Learning · Computer Science 2020-06-30 Matthew B. A. McDermott , Tzu Ming Harry Hsu , Wei-Hung Weng , Marzyeh Ghassemi , Peter Szolovits

Ample evidence suggests that better machine learning models may be steadily obtained by training on increasingly larger datasets on natural language processing (NLP) problems from non-medical domains. Whether the same holds true for medical…

Computation and Language · Computer Science 2020-10-29 Jean-Baptiste Lamare , Tobi Olatunji , Li Yao