English
Related papers

Related papers: Sample Size in Natural Language Processing within …

200 papers

Recent prompt optimisation approaches use the generative nature of language models to produce prompts -- even rivaling the performance of human-curated prompts. In this paper, we demonstrate that randomly sampling tokens from the model…

Computation and Language · Computer Science 2024-04-18 Yao Lu , Jiayi Wang , Raphael Tang , Sebastian Riedel , Pontus Stenetorp

Objectives: Estimation of areas under receiver operating characteristic curves (AUCs) and their differences is a key task in diagnostic studies. We aimed to derive, evaluate, and implement simple sample size formulas for such studies with a…

Methodology · Statistics 2022-08-03 Di Shu , Guangyong Zou

The precipitous rise and adoption of Large Language Models (LLMs) have shattered expectations with the fastest adoption rate of any consumer-facing technology in history. Healthcare, a field that traditionally uses NLP techniques, was bound…

Computation and Language · Computer Science 2023-10-10 Surjya Ray , Pratik Mehta , Hongen Zhang , Ada Chaman , Jian Wang , Chung-Jen Ho , Michael Chiou , Tashfeen Suleman

Automating data extraction from full-text randomised controlled trials (RCTs) for meta-analysis remains a significant challenge. This study evaluates the practical performance of three LLMs (Gemini-2.0-flash, Grok-3, GPT-4o-mini) across…

Computation and Language · Computer Science 2025-07-22 Lingbo Li , Anuradha Mathrani , Teo Susnjak

Text classification is one of the most widely studied tasks in natural language processing. Motivated by the principle of compositionality, large multilayer neural network models have been employed for this task in an attempt to effectively…

Computation and Language · Computer Science 2018-08-07 Devendra Singh Sachan , Manzil Zaheer , Ruslan Salakhutdinov

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on large-scale real-world data such as electronic health records…

Healthcare providers usually record detailed notes of the clinical care delivered to each patient for clinical, research, and billing purposes. Due to the unstructured nature of these narratives, providers employ dedicated staff to assign…

Computation and Language · Computer Science 2022-08-03 Chufan Gao , Mononito Goswami , Jieshi Chen , Artur Dubrawski

In practice, machine learning experts are often confronted with imbalanced data. Without accounting for the imbalance, common classifiers perform poorly and standard evaluation metrics mislead the practitioners on the model's performance. A…

Machine Learning · Computer Science 2020-07-21 Ramiro Camino , Christian Hammerschmidt , Radu State

The increasing volume of healthcare textual data requires computationally efficient, yet highly accurate classification approaches able to handle the nuanced and complex nature of medical terminology. This research presents Knowledge…

Computation and Language · Computer Science 2025-05-13 Hajar Sakai , Sarah S. Lam

Corpus-based methods for natural language processing often use supervised training, requiring expensive manual annotation of training corpora. This paper investigates methods for reducing annotation cost by {\it sample selection}. In this…

cmp-lg · Computer Science 2008-02-03 Sean P. Engelson , Ido Dagan

Objective: Electronic health records (EHR) are widely available to complement administrative data-based disease surveillance and healthcare performance evaluation. Defining conditions from EHR is labour-intensive and requires extensive…

Computation and Language · Computer Science 2025-04-09 Jie Pan , Seungwon Lee , Cheligeer Cheligeer , Elliot A. Martin , Kiarash Riazi , Hude Quan , Na Li

The use of remote sensing in humanitarian crisis response missions is well-established and has proven relevant repeatedly. One of the problems is obtaining gold annotations as it is costly and time consuming which makes it almost impossible…

Computer Vision and Pattern Recognition · Computer Science 2022-02-11 Adrianna Janik , Kris Sankaran

Past studies on the ICD coding problem focus on predicting clinical codes primarily based on the discharge summary. This covers only a small fraction of the notes generated during each hospital stay and leaves potential for improving…

Machine Learning · Computer Science 2023-02-27 Clarence Boon Liang Ng , Diogo Santos , Marek Rei

This study aims to explore the implementation of Natural Language Processing (NLP) and machine learning (ML) techniques to automate the coding of medical letters with visualised explainability and light-weighted local computer settings.…

Computation and Language · Computer Science 2024-07-19 Jamie Glen , Lifeng Han , Paul Rayson , Goran Nenadic

Clinicians may rely on medical coding systems such as International Classification of Diseases (ICD) to identify patients with diseases from Electronic Health Records (EHRs). However, due to the lack of detail and specificity as well as a…

Computation and Language · Computer Science 2022-05-19 Jingqing Zhang , Atri Sharma , Luis Bolanos , Tong Li , Ashwani Tanwar , Vibhor Gupta , Yike Guo

Automated summarization of clinical texts can reduce the burden of medical professionals. "Discharge summaries" are one promising application of the summarization, because they can be generated from daily inpatient records. Our preliminary…

Computation and Language · Computer Science 2022-12-21 Kenichiro Ando , Takashi Okumura , Mamoru Komachi , Hiromasa Horiguchi , Yuji Matsumoto

Purpose. Elevations in initially obtained serum lactate levels are strong predictors of mortality in critically ill patients. Identifying patients whose serum lactate levels are more likely to increase can alert physicians to intensify care…

Quantitative Methods · Quantitative Biology 2021-07-19 Behrooz Mamandipoor , Wesley Yeung , Louis Agha-Mir-Salim , David J. Stone , Venet Osmani , Leo Anthony Celi

Complex survey designs are commonly employed in many medical cohorts. In such scenarios, developing case-specific predictive risk score models that reflect the unique characteristics of the study design is essential for minimizing selective…

Methodology · Statistics 2025-03-27 Marcos Matabuena , Juan C. Vidal , Rahul Ghosal , Jukka-Pekka Onnela

Given a sample of size $N$, it is often useful to select a subsample of smaller size $n<N$ to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the…

Machine Learning · Statistics 2023-10-05 Germain Kolossov , Andrea Montanari , Pulkit Tandon

Small Language Models (SLMs) have potential to be used for automatically labelling and identifying aspects of text data for medicine/health-related purposes from documents and the web. As their resource requirements are significantly lower…

Information Retrieval · Computer Science 2025-11-21 Chris Brogly , Saif Rjaibi , Charlotte Liang , Erica Lam , Edward Wang , Adam Levitan , Sarah Paleczny , Michael Cusimano