English
Related papers

Related papers: Publicly Shareable Clinical Large Language Model B…

200 papers

We introduce SynGP500, a clinician-curated collection of 500 synthetic Australian general practice medical notes. The dataset integrates curriculum-based clinical breadth (RACGP 2022 Curriculum), epidemiologically-calibrated prevalence…

Computation and Language · Computer Science 2025-12-18 Piyawoot Songsiritat

Improving the accuracy and reliability of medical coding reduces clinician burnout and supports revenue cycle processes, freeing providers to focus more on patient care. However, automating the assignment of ICD-10-CM and CPT codes from…

There is enormous enthusiasm and concerns in using large language models (LLMs) in healthcare, yet current assumptions are all based on general-purpose LLMs such as ChatGPT. This study develops a clinical generative LLM, GatorTronGPT, using…

Synthetic Electronic Health Records (EHRs) offer a valuable opportunity to create privacy preserving and harmonized structured data, supporting numerous applications in healthcare. Key benefits of synthetic data include precise control over…

Computation and Language · Computer Science 2025-04-28 Yihan Lin , Zhirong Bella Yu , Simon Lee

Recent studies have demonstrated promising performance of ChatGPT and GPT-4 on several medical domain tasks. However, none have assessed its performance using a large-scale real-world electronic health record database, nor have evaluated…

Computation and Language · Computer Science 2023-07-18 Jingqing Zhang , Kai Sun , Akshay Jagadeesh , Mahta Ghahfarokhi , Deepa Gupta , Ashok Gupta , Vibhor Gupta , Yike Guo

A major obstacle to the development of Natural Language Processing (NLP) methods in the biomedical domain is data accessibility. This problem can be addressed by generating medical data artificially. Most previous studies have focused on…

Computation and Language · Computer Science 2019-08-09 Zixu Wang , Julia Ive , Sumithra Velupillai , Lucia Specia

Physician burnout in the United States has reached critical levels, driven in part by the administrative burden of Electronic Health Record (EHR) documentation and complex diagnostic codes. To relieve this strain and maintain strict patient…

Information Retrieval · Computer Science 2026-03-25 Peter Hartnett , Chung-Chi Huang , Sarah Hartnett , David Hartnett

Generating synthetic text addresses the challenge of data availability in privacy-sensitive domains such as healthcare. This study explores the applicability of synthetic data in real-world medical settings. We introduce MedSyn, a novel…

Computation and Language · Computer Science 2024-09-05 Gleb Kumichev , Pavel Blinov , Yulia Kuzkina , Vasily Goncharov , Galina Zubkova , Nikolai Zenovkin , Aleksei Goncharov , Andrey Savchenko

Despite being a unique source of information on patients' status and disease progression, clinical notes are characterized by high levels of duplication and information redundancy. In general domain text, it has been shown that…

Computation and Language · Computer Science 2023-12-18 Isotta Landi , Eugenia Alleva , Alissa A. Valentine , Lauren A. Lepow , Alexander W. Charney

Synthetic data are becoming a critical tool for building artificially intelligent systems. Simulators provide a way of generating data systematically and at scale. These data can then be used either exclusively, or in conjunction with real…

Artificial Intelligence · Computer Science 2023-04-07 Daniel McDuff , Theodore Curran , Achuta Kadambi

In 2022, with the release of ChatGPT, large-scale language models gained widespread attention. ChatGPT not only surpassed previous models in terms of parameters and the scale of its pretraining corpus but also achieved revolutionary…

Artificial Intelligence · Computer Science 2024-11-13 Yiming Ju , Huanhuan Ma

Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers…

Computation and Language · Computer Science 2025-04-01 Fatemeh Mohammadi , Tommaso Romano , Samira Maghool , Paolo Ceravolo

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical…

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language…

Objective: Clinical knowledge enriched transformer models (e.g., ClinicalBERT) have state-of-the-art results on clinical NLP (natural language processing) tasks. One of the core limitations of these transformer models is the substantial…

Computation and Language · Computer Science 2023-01-30 Yikuan Li , Ramsey M. Wehbe , Faraz S. Ahmad , Hanyin Wang , Yuan Luo

Clinical notes are a rich source of information about patient state. However, using them to predict clinical events with machine learning models is challenging. They are very high dimensional, sparse and have complex structure. Furthermore,…

Machine Learning · Statistics 2018-08-20 Sebastien Dubois , Nathanael Romano , David C. Kale , Nigam Shah , Kenneth Jung

Clinical notes contain an extensive record of a patient's health status, such as smoking status or the presence of heart conditions. However, this detail is not replicated within the structured data of electronic health systems.…

Computation and Language · Computer Science 2020-09-18 Andriy Mulyar , Elliot Schumacher , Masoud Rouhizadeh , Mark Dredze

High computation costs and latency of large language models such as GPT-4 have limited their deployment in clinical settings. Small language models (SLMs) offer a cost-effective alternative, but their limited capacity requires biomedical…

Alongside the growth of generative AI, we are witnessing a surge in the use of synthetic data across all stages of the AI development pipeline. It is now common practice for researchers and practitioners to use one large generative model…

Human-Computer Interaction · Computer Science 2025-05-14 Shivani Kapania , Stephanie Ballard , Alex Kessler , Jennifer Wortman Vaughan

Social media datasets are essential for research on a variety of topics, such as disinformation, influence operations, hate speech detection, or influencer marketing practices. However, access to social media datasets is often constrained…

Computation and Language · Computer Science 2025-05-07 Henry Tari , Nojus Sereiva , Rishabh Kaushal , Thales Bertaglia , Adriana Iamnitchi