English
Related papers

Related papers: Development and validation of a natural language p…

200 papers

The prevalence of ambiguous acronyms make scientific documents harder to understand for humans and machines alike, presenting a need for models that can automatically identify acronyms in text and disambiguate their meaning. We introduce…

Computation and Language · Computer Science 2021-01-07 Nicholas Egan , John Bohannon

The use of Natural Language Processing (NLP) in highstakes AI-based applications has increased significantly in recent years, especially since the emergence of Large Language Models (LLMs). However, despite their strong performance, LLMs…

Protected health information (PHI) de-identification is critical for enabling the safe reuse of clinical notes, yet evaluating and comparing PHI de-identification models typically depends on costly, small-scale expert annotations. We…

Artificial Intelligence · Computer Science 2025-11-19 Guanchen Wu , Zuhui Chen , Yuzhang Xie , Carl Yang

De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014 are over a decade old and lack the semantic and demographic diversity of modern…

Computation and Language · Computer Science 2026-05-06 Jose D. Posada , David Love , Somalee Datta , Priya Desai

Leveraging medical record information in the era of big data and machine learning comes with the caveat that data must be cleaned and de-identified. Facilitating data sharing and harmonization for multi-center collaborations are…

Image and Video Processing · Electrical Eng. & Systems 2023-05-11 Adrienne Kline , Vinesh Appadurai , Yuan Luo , Sanjiv Shah

Data cleaning consumes about 80% of the time spent on data analysis for clinical research projects. This is a much bigger problem in the era of big data and machine learning in the field of medicine where large volumes of data are being…

Medical Physics · Physics 2018-01-03 Timothy Rozario , Troy Long , Mingli Chen , Weiguo Lu , Steve Jiang

Objective:Develop and validate an algorithm for analyzing the layout of PDF clinical documents to improve the performance of downstream natural language processing tasks. Materials and Methods: We designed an algorithm to process clinical…

Computation and Language · Computer Science 2023-05-24 Christel Gérardin , Perceval Wajsbürt , Basile Dura , Alice Calliger , Alexandre Moucher , Xavier Tannier , Romain Bey

Structuring medical data in France remains a challenge mainly because of the lack of medical data due to privacy concerns and the lack of methods and approaches on processing the French language. One of these challenges is structuring…

Computation and Language · Computer Science 2021-12-22 Azzam Alwan , Maayane Attias , Larry Rubin , Adnan El Bakri

Accessing sensitive patient data for machine learning is challenging due to privacy concerns. Datasets with annotations of personally identifiable information are crucial for developing and testing anonymization systems to enable safe data…

Computation and Language · Computer Science 2026-03-17 Ibrahim Baroud , Christoph Otto , Vera Czehmann , Christine Hovhannisyan , Lisa Raithel , Sebastian Möller , Roland Roller

The unstructured nature of clinical notes within electronic health records often conceals vital patient-related information, making it challenging to access or interpret. To uncover this hidden information, specialized Natural Language…

Access to medical imaging and associated text data has the potential to drive major advances in healthcare research and patient outcomes. However, the presence of Protected Health Information (PHI) and Personally Identifiable Information…

Machine Learning · Statistics 2025-08-01 Kyle Naddeo , Nikolas Koutsoubis , Rahul Krish , Ghulam Rasool , Nidhal Bouaynaya , Tony OSullivan , Raj Krish

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or replacing sensitive information from data, enabling its…

Computation and Language · Computer Science 2020-03-18 Aitor García-Pablos , Naiara Perez , Montse Cuadros

The widespread adoption of electronic health records has created new opportunities for translational clinical research, yet this promise remains constrained by fragmented data across privacy-siloed institutions and substantial heterogeneity…

Since the COVID-19 pandemic, clinicians have seen a large and sustained influx in patient portal messages, significantly contributing to clinician burnout. To the best of our knowledge, there are no large-scale public patient portal…

Artificial Intelligence · Computer Science 2024-11-12 Joseph Gatto , Parker Seegmiller , Timothy E. Burdick , Sarah Masud Preum

Large Language Models (LLMs) are increasingly adopted across domains such as education, healthcare, and finance. In healthcare, LLMs support tasks including disease diagnosis, abnormality classification, and clinical decision-making. Among…

Cryptography and Security · Computer Science 2026-03-31 Payel Bhattacharjee , Fengwei Tian , Geoffrey D. Rubin , Joseph Y. Lo , Nirav Merchant , Heidi Hanson , John Gounley , Ravi Tandon

Many models are pretrained on redacted text for privacy reasons. Clinical foundation models are often trained on de-identified text, which uses special syntax (masked) text in place of protected health information. Even though these models…

Computation and Language · Computer Science 2025-06-18 Paul Landes , Aaron J Chaise , Tarak Nath Nandi , Ravi K Madduri

Background Clinical studies using real-world data may benefit from exploiting clinical reports, a particularly rich albeit unstructured medium. To that end, natural language processing can extract relevant information. Methods based on…

Computation and Language · Computer Science 2022-07-27 Basile Dura , Charline Jean , Xavier Tannier , Alice Calliger , Romain Bey , Antoine Neuraz , Rémi Flicoteaux

Applying methods in natural language processing on electronic health records (EHR) data is a growing field. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there is a paucity of annotated…

Computation and Language · Computer Science 2022-04-08 Yanjun Gao , Dmitriy Dligach , Timothy Miller , Samuel Tesch , Ryan Laffin , Matthew M. Churpek , Majid Afshar

The recognition of medical entities from natural language is an ubiquitous problem in the medical field, with applications ranging from medical act coding to the analysis of electronic health data for public health. It is however a complex…

Computation and Language · Computer Science 2020-05-07 Louis Falissard , Claire Morgand , Sylvie Roussel , Claire Imbaud , Walid Ghosn , Karim Bounebache , Grégoire Rey

Acronyms are the short forms of longer phrases and they are frequently used in writing, especially scholarly writing, to save space and facilitate the communication of information. As such, every text understanding tool should be capable of…

Computation and Language · Computer Science 2021-01-07 Amir Pouran Ben Veyseh , Franck Dernoncourt , Thien Huu Nguyen , Walter Chang , Leo Anthony Celi