中文
相关论文

相关论文: Sample Size in Natural Language Processing within …

200 篇论文

This paper explores an approach to Bayesian sample size determination in clinical trials. The approach falls into the category of what is often called "proper Bayesian", in that it does not mix frequentist concepts with Bayesian ones. A…

统计方法学 · 统计学 2012-04-23 Robb J. Muirhead , Adina I. Soaita

This work presents a systematic and in-depth investigation of the utility of large language models as text classifiers for biomedical article classification. The study uses several small and mid-size open source models, as well as selected…

计算与语言 · 计算机科学 2026-03-13 Jakub Proboszcz , Paweł Cichosz

In biospectroscopy, suitably annotated and statistically independent samples (e. g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training…

应用统计 · 统计学 2015-05-05 Claudia Beleites , Ute Neugebauer , Thomas Bocklitz , Christoph Krafft , Jürgen Popp

Although machine learning has become a powerful tool to augment doctors in clinical analysis, the immense amount of labeled data that is necessary to train supervised learning approaches burdens each development task as time and resource…

With large volumes of health care data comes the research area of computational phenotyping, making use of techniques such as machine learning to describe illnesses and other clinical concepts from the data itself. The "traditional"…

机器学习 · 统计学 2016-12-30 Chris Hodapp

The objective of many high-dimensional microarray and RNA-seq studies is to develop a classifier of cancer patients based on characteristics of their disease. The germinal center B-cell (GCB) classifier study in lymphoma and the National…

应用统计 · 统计学 2015-09-17 Sandra Safo , Xiao Song , Kevin K. Dobbin

Recent work has shown that fine-tuning large networks is surprisingly sensitive to changes in random seed(s). We explore the implications of this phenomenon for model fairness across demographic groups in clinical prediction tasks over…

计算与语言 · 计算机科学 2021-04-14 Silvio Amir , Jan-Willem van de Meent , Byron C. Wallace

Large language models (LLMs) are increasingly deployed in real-world systems, yet they can produce toxic or biased outputs that undermine safety and trust. Post-hoc model repair provides a practical remedy, but the high cost of parameter…

机器学习 · 计算机科学 2025-10-24 Xuran Li , Jingyi Wang

Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a…

计算与语言 · 计算机科学 2025-01-03 Tomek Rutowski , Amir Harati , Elizabeth Shriberg , Yang Lu , Piotr Chlebek , Ricardo Oliveira

Statistical samples, in order to be representative, have to be drawn from a population in a random and unbiased way. Nevertheless, it is common practice in the field of model-based diagnosis to make estimations from (biased) best-first…

人工智能 · 计算机科学 2022-08-05 Patrick Rodler , Fatima Elichanova

Clinical notes are often stored in unstructured or semi-structured formats after extraction from electronic medical record (EMR) systems, which complicates their use for secondary analysis and downstream clinical applications. Reliable…

计算与语言 · 计算机科学 2025-12-30 Risha Surana , Adrian Law , Sunwoo Kim , Rishab Sridhar , Angxiao Han , Peiyu Hong

Deep learning models in healthcare may fail to generalize on data from unseen corpora. Additionally, no quantitative metric exists to tell how existing models will perform on new data. Previous studies demonstrated that NLP models of…

计算与语言 · 计算机科学 2021-02-22 Mihir P. Khambete , William Su , Juan Garcia , Marcus A. Badgeley

We consider a Bayesian framework for estimating the sample size of a clinical trial. The new approach, called BESS, is built upon three pillars: Sample size of the trial, Evidence from the observed data, and Confidence of the final decision…

统计方法学 · 统计学 2026-01-21 Dehua Bi , Yuan Ji

Many practical applications of AI in medicine consist of semi-supervised discovery: The investigator aims to identify features of interest at a resolution more fine-grained than that of the available human labels. This is often the scenario…

计算与语言 · 计算机科学 2020-04-08 Allen Schmaltz , Andrew Beam

Diabetes is a chronic disease with a significant global health burden, requiring multi-stakeholder collaboration for optimal management. Large language models (LLMs) have shown promise in various healthcare scenarios, but their…

The increasing volume and complexity of clinical documentation in Electronic Medical Records systems pose significant challenges for clinical coders, who must mentally process and summarise vast amounts of clinical text to extract essential…

计算与语言 · 计算机科学 2024-09-25 Bokang Bi , Leibo Liu , Sanja Lujic , Louisa Jorm , Oscar Perez-Concha

While there exists a large amount of literature on the general challenges of and best practices for trustworthy online A/B testing, there are limited studies on sample size estimation, which plays a crucial role in trustworthy and efficient…

统计方法学 · 统计学 2023-08-21 Jing Zhou , Jiannan Lu , Anas Shallah

The application of large language models (LLMs) to healthcare information extraction has emerged as a promising approach. This study evaluates the classification performance of five open-source LLMs: GEMMA-3-27B-IT, LLAMA3-70B, LLAMA4-109B,…

计算与语言 · 计算机科学 2025-05-09 Yuting Guo , Abeed Sarker

The advent of Large Language Models (LLMs) is promising and LLMs have been applied to numerous fields. However, it is not trivial to implement LLMs in the medical field, due to the high standards for precision and accuracy. Currently, the…

信息检索 · 计算机科学 2024-12-04 Rishabh Goel

Medical coding is the task of assigning medical codes to clinical free-text documentation. Healthcare professionals manually assign such codes to track patient diagnoses and treatments. Automated medical coding can considerably alleviate…